AI Detectors Beyond English: Spanish, French, Chinese and More
Almost every AI detector you have ever used was built by people who think in English, trained on text written mostly in English, and tuned against a benchmark of English essays. That is not a slur against the engineers who built these tools. It is simply where the data was, where the money was, and where the first wave of panic about ChatGPT landed. But it means that the moment you feed a detector a paragraph of Spanish, French, Mandarin, Arabic, or Swahili, you have quietly stepped outside the conditions the tool was designed for. The numbers on the marketing page still say ninety-something percent. The confidence badge still glows green or red. What changed is that the ground underneath those numbers got a great deal softer, and almost nobody tells you so.
This article is about that softness. If you teach, grade, hire, moderate, or publish in any language other than English, you are running these tools in a regime where they are weaker, more erratic, and more likely to hurt the wrong person than the English-first world that produced them. Understanding why is not academic. It changes whether you should trust a flag at all, and it changes who pays the price when the tool gets it wrong.
Why English became the default, and why that matters everywhere else
Large language models learn from whatever text they are fed, and the internet's text is lopsided. English is over-represented by an enormous margin relative to its share of the world's speakers. When a model like GPT-4 or its successors is trained, it sees far more English than any other language, so it becomes most fluent, most idiomatic, and most statistically well-behaved in English. Detectors inherit this lopsidedness twice over. First, the underlying language models they lean on are English-heavy. Second, the detectors themselves are trained and validated on datasets of AI-generated versus human text that are overwhelmingly English, because that is what researchers had lying around and what customers were asking about.
The consequence is subtle but decisive. A detector does not really "understand" whether text is machine-written. It looks for statistical fingerprints — the way word choices distribute, how predictable each next token is, how much the sentence rhythm varies. Those fingerprints were characterized in English. In a language the detector has seen far less of, the fingerprints are blurrier, the baseline for "normal human writing" is shakier, and the boundary between human and machine is drawn with far less evidence. You are asking a tool calibrated on one language to make fine-grained judgments in another, and calibration does not transfer for free.
The mechanics: why perplexity and burstiness misbehave across languages
Most detectors, under the hood, are chasing two related signals. Perplexity is a measure of how surprised a language model is by a piece of text — AI-generated text tends to be low-perplexity because the model that wrote it, or a model like it, finds it very predictable. Burstiness describes the variation in sentence length and complexity — human writing tends to lurch between long and short, dense and sparse, while machine text often flows at an eerily even pace. If you want the full mechanical picture of how these signals work and where they break, we walk through it in our explainer on perplexity, burstiness, and AI watermarks. The short version for our purposes here is that both signals are language-relative, and neither travels cleanly across a border.
Consider perplexity first. A detector estimates perplexity using a reference language model. If that reference model is far more confident in English than in, say, Hungarian or Tagalog, then its perplexity readings in those languages are noisier and its sense of "surprisingly predictable" is distorted. Human writing in a language the reference model handles poorly can look artificially low-perplexity — or artificially high — for reasons that have nothing to do with whether a machine wrote it. The measuring stick itself is warped.
Burstiness is even more culturally loaded than people assume. Sentence-length variation is not a universal constant. Languages have different average sentence structures, different punctuation conventions, and different rhetorical traditions. Written Chinese and Japanese segment ideas differently from English; classical rhetorical patterns in Arabic favor longer, more elaborate constructions; German technical prose tolerates long compound sentences that would feel unnatural in English. A detector that learned "humans are bursty, machines are smooth" from English essays is applying a yardstick that was never validated against these other rhythms. Perfectly ordinary human prose in another language can register as suspiciously uniform simply because its native cadence does not match the English pattern the tool encoded as "human."
There is a second-order problem too. Tokenization — the way text gets chopped into units the model processes — is far less efficient for many non-English languages, especially those that do not use spaces between words or that rely on rich morphology. Non-Latin scripts often consume more tokens per unit of meaning, which changes the granularity at which the detector "sees" the text and further muddies the statistical signals it depends on. None of this is visible to the person reading a confidence score. It just quietly degrades the reliability of that score.
The double bias: under-detection in one direction, over-flagging in the other
Here is where the story turns genuinely unfair, because the failures do not cancel out. They compound in two different directions, and both directions land on real people.
The first failure is under-detection. In languages where the detector is weak, AI-generated text can slip through more easily. A student who runs a Spanish essay through a translator, or who prompts a model to write directly in Portuguese, may face a tool that simply cannot characterize the output confidently enough to flag it. Content mills producing bulk AI text in Indonesian, Vietnamese, or Turkish operate in a comfortable blind spot. The tool's silence is not evidence of human authorship; it is evidence that the tool ran out of confidence and defaulted to a shrug. Anyone relying on a detector to keep a non-English channel clean is likely getting far weaker protection than the English case study suggested.
The second failure is the one that should keep educators up at night: over-flagging of genuine human writing, and it falls disproportionately on non-native English speakers. This finding has been widely reported and repeatedly reproduced. Essays written by real human students who learned English as a second or additional language are flagged as AI-generated at strikingly higher rates than essays by native speakers. The mechanism is almost poetic in its cruelty. Non-native writers tend to use a more limited, more common vocabulary, more predictable sentence constructions, and fewer idiosyncratic flourishes — precisely the low-perplexity, low-burstiness profile that detectors read as machine-like. The very features of careful, learned, second-language writing are the features detectors were trained to treat as suspicious.
So you get a grimly asymmetric outcome. Fluent native speakers producing polished English sail through, and so, often, does the AI text that mimics that polish. Meanwhile a diligent international student writing in their third language, entirely by hand, gets accused of cheating because their prose is "too clean." We treat this pattern at length in our piece on how AI detector false positives happen, but it deserves emphasis in the multilingual context specifically, because language ability and immigration status and educational access are all tangled up in who gets caught in this net. The detector is not neutral. It encodes a preference for a particular register of native-speaker English and penalizes deviation from it, whether that deviation comes from a machine or from a human being still mastering the language.
A language-by-language reality check
It is tempting to talk about "non-English" as one undifferentiated blob, but the reliability varies enormously depending on which language you are actually dealing with. Broadly, the picture looks like a gradient.
At the more supported end sit the high-resource European languages closest to English in structure and in training-data volume: Spanish, French, German, Italian, Portuguese, Dutch. Several detectors handle these with something approaching their English-tier confidence — not equal to it, but in the same neighborhood. There is a lot of Spanish and French on the internet, these languages share the Latin script and a good deal of grammatical machinery with English, and vendors have had commercial reasons to invest in them. If you are grading French or Spanish, you are in the least-bad part of the map. "Least bad" is still not "good," and the false-positive risk for non-native writers persists, but the raw detection signal is more meaningful here than almost anywhere else outside English.
The middle of the gradient holds languages that are widely spoken and reasonably well-resourced but structurally distant from English or written in non-Latin scripts: Russian, Turkish, Polish, Vietnamese, Indonesian, Hindi. Coverage exists on paper, but confidence and consistency drop. The detector has seen enough to try, not enough to be sure.
At the weak end sit two overlapping categories. First, major world languages that are structurally and orthographically far from English — Chinese, Japanese, Korean, Arabic. These are spoken by enormous populations and are anything but obscure, yet detectors tend to be markedly less reliable in them because the statistical patterns, scripts, tokenization behavior, and rhetorical conventions diverge so sharply from the English training regime. Chinese and Japanese in particular expose the burstiness assumptions as parochial. Arabic's morphology and rhetorical tradition confound tools tuned to English cadence. Second, and worst of all, the genuinely low-resource languages — Swahili, Amharic, Bengali dialects, most African and indigenous languages, and countless others — where there is so little training data that detection is close to guesswork. A confidence score in one of these languages should be treated as decorative.
The uncomfortable takeaway is that detector reliability roughly tracks a language's economic and digital privilege. The languages of wealthy, heavily-online populations get the best coverage; the languages of the global majority get the worst. That is not a coincidence, and it has real consequences for who can be falsely accused and who can quietly get away with things.
What "supports 30 languages" actually means on a pricing page
Vendors have noticed that multilingual capability sells, and the marketing has adjusted accordingly. You will see confident claims: this tool supports thirty languages, that one supports over one hundred. Copyleaks advertises broad multilingual coverage across dozens of languages; Originality.ai has expanded its language claims; GPTZero and others tout varying degrees of support. These claims are not lies, exactly. But there is a crucial gap between "supports" and "is accurate in," and the marketing lives in that gap.
"Supports language X" typically means the tool will accept text in that language and return a score rather than an error. It does not mean the tool has been rigorously validated in that language, that its false-positive rate has been measured against native and non-native writers in that language, or that its confidence in that language matches its confidence in English. A tool can technically "support" a hundred languages while being genuinely reliable in only a handful. The support list is a compatibility statement, not an accuracy guarantee, and vendors rarely publish per-language accuracy breakdowns because those breakdowns would be unflattering.
When we looked closely at one of the most prominent multilingual claimants in our Copyleaks detector review, the pattern held: broad language coverage on paper, far less transparency about how performance actually varies from one language to the next. This is not unique to any single vendor. The honest question to ask any tool is not "how many languages do you support" but "where is your published false-positive rate for human-written text in the specific language I care about, broken out for native and non-native writers." Almost no vendor can answer that, and the silence is the answer. Our broader look at what the accuracy data actually shows in 2026 reinforces the point that headline numbers, even in English, hide a lot of variance — and that variance only widens as you move across languages.
Translation laundering: the loophole nobody advertises
There is a specific and widely-known technique that exploits the multilingual weakness directly, and it is worth understanding even if you have no intention of using it, because it tells you how brittle the detection signal really is.
The maneuver goes like this. Someone generates text with an AI model in English, then runs it through machine translation into another language — or generates in English, translates to a second language, and sometimes translates back. Each pass through a translation system rewrites the surface form of the text. The word choices change, the sentence structures get reshuffled, the token-level predictability that a detector keys on gets scrambled. The AI fingerprint that a detector could have caught in the original English is smeared out. What comes out the other side often reads, to a detector, like something it can no longer confidently attribute to a machine.
Translation laundering works precisely because detectors are surface-pattern matchers rather than meaning-comprehenders. They are not reasoning about whether the ideas are machine-generated; they are measuring statistical properties of the specific words in the specific order they appear. Rewrite those words and reorder them — which is exactly what translation does — and you have altered the very thing being measured without changing the underlying fact that a machine produced the content. It is the multilingual cousin of the paraphrasing and "humanizing" tricks that plague English detection, and it is arguably more effective because translation is a more thorough rewrite than most paraphrasers manage. The same fragility shows up whenever text is heavily rewritten; our review of tools built around cross-language academic integrity, including the Crossplag detector, touches on how translation and cross-lingual reuse complicate detection in ways single-language tools were never designed to handle.
The practical lesson is not that everyone is laundering AI text through translation. It is that a detector's inability to flag translated content should never be read as a clean bill of health. Absence of a flag in a multilingual pipeline is close to meaningless.
Who actually pays for this
Every technical weakness described so far converts, eventually, into a consequence for a specific human being, and the distribution of those consequences is not random. It is the equity dimension that makes multilingual detection more than a curiosity for engineers.
Consider the international student. They may be studying in an English-medium university while thinking, drafting, and sometimes writing in their first language. They are exactly the population most likely to write clean, careful, slightly-formulaic English — the profile detectors over-flag — and exactly the population least equipped to defend themselves when accused. An accusation of academic dishonesty is catastrophic for a student on a visa; it can end a degree, a scholarship, an immigration status. When a detector's false positive rides on top of a language barrier and a power imbalance, the student who did nothing wrong is the one who suffers, and they are often the person least able to push back against an instructor waving a confidence score.
Consider the global content ecosystem. Publishers, marketplaces, and platforms increasingly use detectors to gate or penalize AI content. In English-speaking markets this is already crude; in other-language markets it is cruder still. Legitimate human writers producing content in Hindi or Portuguese or Arabic get caught by tools that cannot reliably tell their careful human prose from a machine's, while actual AI content mills operating in poorly-covered languages sail past. The tool ends up punishing the honest small-language writer and rewarding the bad actor who picked the right blind spot. That is the opposite of what the tool was bought to do.
And consider the broader signal it sends. When detection is reliable for English and unreliable for everyone else, and when the errors specifically punish non-native speakers, the technology quietly encodes a hierarchy. It tells a Nigerian writer, a Filipino nurse writing an application essay, a Peruvian graduate student, that their careful work is presumptively suspect in a way a native Anglophone's never would be. Whatever the intentions of the people who built these tools, that is the lived effect of deploying them at scale outside the language they were made for.
Honest guidance for using these tools in a multilingual world
None of this means detectors are useless everywhere but English. It means you have to hold them far more loosely the moment you cross a language line, and you have to design your process around their known weaknesses rather than around their marketing.
Start by treating the language of the text as a first-class variable in how much weight you give a result. A flag on a Spanish or French document deserves more consideration than a flag on a Chinese, Arabic, or low-resource-language document — but even the European-language flag is not proof, only a prompt to look closer. In the weaker languages, a score should barely move your prior at all. If you would not accept a coin flip as evidence, do not accept a detector's verdict in a language where its behavior is close to a coin flip.
Never let a detector's output be the sole basis for an accusation in any language, and be especially disciplined about this with non-native writers. The over-flagging finding is not a marginal caveat; it is a central, reproducible feature of how these tools behave. If a piece of writing seems "too clean," ask whether "clean" is really "AI" or whether it is simply "written by someone who learned this language methodically." Look for the human evidence a detector cannot see: drafts, version history, the writer's process, an actual conversation. A five-minute discussion in which a student explains their argument tells you more than any confidence percentage.
Be honest with the people you assess about what you are using and what its limits are. If you run non-English work through a detector, you owe the writer transparency about the tool's known unreliability in their language and a real avenue to contest a result. Building the appeal in from the start is not a courtesy; it is the only defensible way to use an instrument you know misfires along linguistic and national lines.
Interrogate vendor claims with the specific question that matters: not the length of the supported-languages list, but the existence of a published, per-language false-positive rate covering both native and non-native writers. Treat the absence of that number as the disqualifying fact it is. A tool that will not tell you how often it wrongly accuses a human in the language you care about has told you everything you need to know about how confident you should be.
Finally, keep the whole enterprise in proportion. Detection is a probabilistic hint even at its English-language best, and outside English it degrades from a hint into a rumor. The people most likely to be harmed by treating that rumor as fact are the ones with the least power to defend themselves — the international student, the second-language writer, the small-language creator. A tool that cannot reliably distinguish a machine from a diligent human speaking their third language has not earned the authority to end anyone's academic career or livelihood. Use these detectors, if you use them at all, as one soft input among many, held loosest exactly where the marketing insists they are strong. In every language but the one they were born in, humility is not optional — it is the whole method.