AI Detector Percentages: What a Score Actually Means
You paste a paragraph into a detector, wait a second, and a number appears: 63% AI. Almost everyone reads that the same way on first glance — sixty-three percent of these words came from a machine, the other thirty-seven from a human. It feels like a proportion, a mixture, a recipe you could reverse-engineer. It is not. That single misreading is the source of most of the panic, the arguments, and the bad decisions that happen after a score lands on screen. The number is real, but it almost never means what your gut tells you it means. This piece is about decoding that number — what it is actually expressing, why the same passage produces different figures depending on where you paste it, and what you should and shouldn't do once you're staring at one.
The percentage is a guess about origin, not a measurement of content
Start with the thing the number is not. A detector does not read your text, identify machine-written sentences, tally them, and divide by the total. It has no ledger of "these words are synthetic." What it produces is closer to a confidence estimate — the model's own reading of how likely it is that text of this shape was generated by a language model rather than typed by a person. When a tool reports 63%, the honest translation is usually something like: "Given the patterns I've learned to associate with machine writing, I'm moderately confident this document was AI-generated." That's a statement about the detector's certainty, not about the composition of your paragraph.
The distinction sounds academic until you watch it change behavior. If 63% meant "63% of the words are AI," you'd expect a document that's genuinely half-human and half-machine to land near 50%. It usually doesn't. A lightly edited AI draft can score in the high nineties, because the statistical fingerprint of the original generation survives your edits. A fully human essay written in a clean, even, well-organized style can score in the seventies, because clean and even is exactly what machines also produce. The percentage tracks resemblance to a pattern, not the literal share of synthetic sentences. Once you internalize that, half the confusion dissolves.
This is also why the number can feel slippery. A proportion is stable — half a cake is half a cake no matter who's holding the knife. A probability estimate is a judgment, and judgments shift with the evidence, the model, and the yardstick being used. The moment you treat the percentage as a probability rather than a proportion, its odd behavior starts to make sense.
Three different things a "percentage" can secretly be
Here's where it gets genuinely messy: vendors don't agree on what their percentage represents. Two tools can both show you "78%" and mean structurally different things by it. Broadly, the number you see is one of three creatures wearing the same costume.
The first is a document-level probability. The tool assesses the whole passage as a single unit and reports one figure for how AI-like the entire thing reads. This is the most common flavor for consumer detectors. "82% AI" here means "my overall confidence that this document is machine-generated is 82%." It says nothing about any individual sentence.
The second is a proportion built from per-sentence verdicts. Some tools score each sentence, label it as AI or human, and then report the fraction of sentences that tripped the AI label. In that world, "40%" genuinely does mean "40% of the sentences here looked machine-written to me." This is the one case where the intuitive reading is roughly correct — but you usually can't tell from the number alone whether you're looking at this kind of tool or the first kind.
The third is an aggregate confidence score that blends per-sentence signals into a single weighted figure using the tool's own internal formula. It's neither a clean probability nor a clean proportion; it's a manufactured index. Some sentences count more than others, thresholds get applied along the way, and the final figure is a summary the vendor designed to be readable rather than to be interpretable. When people say "the percentage is basically made up," this is the version they're reacting to — not made up in the sense of being random, but constructed rather than measured.
The uncomfortable takeaway is that the same number carries different meaning across tools, and the interface rarely tells you which kind you're holding. "78%" from a document-probability tool and "78%" from a sentence-proportion tool are not comparable quantities. They just happen to render as the same two digits. Treating them as the same is like comparing a temperature in Celsius to one in Fahrenheit because both readings say "40."
Why one passage scores three different ways in three tabs
Paste an identical paragraph into three detectors and you'll often get three notably different numbers — 30% here, 55% there, 88% in the third. People find this maddening, and reasonably so: if these tools measured the same physical property, they'd agree the way three thermometers agree. They disagree because they aren't measuring one fixed property. Each detector was trained on different text, on outputs from different generations of models, with different definitions of what "AI-like" even looks like, and each one draws its dividing line in a different place.
Under the hood, most detectors lean on statistical signals like predictability and variation — how unsurprising each next word is, and how much the rhythm of sentence length and complexity fluctuates across the passage. If you want the mechanics of those signals, we cover them in perplexity, burstiness, and watermarks explained. But two tools can weigh those same signals differently, calibrate them against different reference text, and arrive at genuinely different confidence levels for the identical input. Neither is lying. They're answering slightly different questions and reporting the answers in the same units.
There's a deeper reason for the spread, too. A detector trained mostly on last year's model outputs may be miscalibrated against text from a newer model, or against a human who happens to write in a plain, orderly style. The dividing line each vendor chose was tuned on their data, for their goals, at their moment in time. We unpack the full set of reasons in why AI detectors give different results, but the short version is this: divergence between tools isn't a bug you can average your way out of. It's a signal that the number is a judgment, and different judges disagree.
The false-precision trap: "it said 82%" tells me almost nothing
Two digits and a percent sign radiate authority. They look measured, calibrated, final — the way a lab result or a bank balance looks. That polish is exactly the problem. A screenshot that says "82% AI" and nothing else is, on its own, close to meaningless, and it's worth being blunt about why.
To interpret 82% you need at least three things the number doesn't carry with it. You need to know which tool produced it, because 82% means different things across vendors, as we've seen. You need to know what the tool's threshold is — whether 82% clears the line the vendor treats as "flagged" or sits comfortably below it. And you need to know how the tool behaves on this kind of writing — some detectors run consistently hot on formal, structured prose regardless of who wrote it. Strip those three away and the bare figure is a number without a unit. It's like being told a temperature of "82" with no scale, no location, and no thermometer.
The false precision gets amplified when the number changes hands. A student sees 82%, screenshots it, and it becomes "the detector said my essay is 82% AI" in the retelling — as if it were a settled fact rather than one tool's shaky estimate at one moment. An instructor forwards it and it hardens further. Each retelling sands off the uncertainty and leaves only the clean, confident digits. The more precise a number looks, the more we trust it, and detector percentages look extremely precise while resting on genuinely uncertain ground. Precision and accuracy are not the same thing, and detectors routinely offer the first while quietly lacking the second.
Where the line gets drawn, and what it costs
Every detector that renders a verdict — flagged versus clear, red versus green — has a threshold hiding behind the percentage: a cutoff above which the tool decides "call this AI." That cutoff is a choice, not a law of nature, and moving it forces an unavoidable trade-off that every detector vendor has to make and mostly doesn't advertise.
Set the threshold low — say, flag anything above 30% — and you catch more genuine AI use, but you also start flagging honest human writers who happen to write in a plain, predictable style. Set it high — flag only above 90% — and you protect the innocent, but a lot of real machine text slides through as "clear." There's no setting that eliminates both errors at once. Push down the rate of missed AI and you push up the rate of falsely accused humans; relax one and the other tightens. This is the fundamental tension baked into every detector, and it's why the same raw score can be reported as "flagged" by one tool and "likely human" by another that simply drew its line elsewhere.
The false-positive side of that trade is where the real human cost lands. A missed detection is a machine essay that goes ungraded-as-AI — annoying, but nobody gets hurt. A false positive is a real person, who wrote every word themselves, being told an algorithm thinks they cheated. Those aren't symmetric harms, which is why the responsible framing treats a false positive as far more expensive than a missed catch. We go deep on this in AI detector false positives explained, and it's the single most important thing to understand before anyone acts on a score. The percentage you see already has a vendor's opinion about that trade-off baked into where "flagged" begins.
What 0% and 100% actually tell you
The two ends of the scale are the most misread of all, so they're worth taking one at a time.
The zero that isn't innocence
A 0% score feels like an acquittal — the machine looked and found nothing. What it actually means is narrower: this text did not trigger the patterns this particular tool associates with AI writing. That's not the same as "definitely human." A skilled user can prompt a model to write in a loose, uneven, idiosyncratic voice that sails under the detector's radar and scores near zero. Meanwhile, plenty of genuinely human writing scores near zero simply because it's messy in the ways machines usually aren't. A 0% is best read as "no signal here," not as a certificate. Absence of evidence, in detector terms, is not evidence of absence.
The hundred that's almost never earned
At the other end, 100% looks like the machine caught someone red-handed with total certainty. But a truly honest probability model should almost never output absolute certainty, because certainty means zero possibility of being wrong — and no detector is that good. When a tool shows a flat 100%, it's usually the display rounding up from something like 99.4%, or the interface simply capping the scale at a round ceiling. It is a design choice about how to present a very high confidence, not a claim that error is impossible. Treating 100% as literal proof mistakes a rounded, styled number for a guarantee the underlying model never actually made. The higher the confidence a detector projects, the more skeptical you should be that it has earned it — because overconfidence is exactly the failure mode these systems are prone to.
Sentence-level versus document-level: two different maps
Many detectors offer two views: one overall percentage for the whole document, and a sentence-by-sentence highlight showing which passages the tool considers suspicious. These are different maps of the same territory, and confusing them leads people astray fast.
The document-level figure is a single summary judgment. The sentence-level view is more granular but also more fragile: individual sentences carry far less text for the model to work with, so per-sentence verdicts swing hard on tiny inputs. A short, plain sentence — "The results were clear." — carries almost no statistical signal and can get flagged or cleared on a coin-flip basis. Longer, more distinctive passages give the model more to go on and produce steadier verdicts. This is why the highlighted-sentence view often looks scattered and arbitrary: it is scattered, because short spans are genuinely hard to classify.
The practical mistake is reading the highlights as a precise accusation — "these exact three sentences are the AI ones." They usually aren't a reliable map of anything that fine-grained. Treat the sentence view as a rough heat map of where the tool's suspicion clusters, not as a line-item indictment. And when the document-level number and the sentence-level highlights disagree — a low overall score but a few red sentences, or vice versa — that disagreement is information about the tool's uncertainty, not a puzzle you're meant to solve into a verdict. If the underlying signals of how detectors read text interest you, what an AI detector is and how detection works lays out the whole mechanism.
So a number appeared. Now what?
How you should respond depends heavily on which chair you're sitting in, because the same percentage means a different thing to a student, a teacher, and a working writer.
If you're a student staring at a flag
First, breathe — a percentage is not a verdict, and a detector cannot actually see whether you wrote your essay. If you wrote it yourself, your best evidence isn't a counter-score from a different detector; it's your process. Draft history, version records, notes, outlines, browser or document timestamps — the trail of how the work came together is far more persuasive to a reasonable person than any tool's number. If you've been flagged despite writing everything yourself, the specific situation of an unjust Turnitin flag is common enough that we wrote a whole guide to it: Turnitin flagged my essay but I didn't use AI. The percentage is the start of a conversation, not the end of one.
If you're a teacher holding a score
The number is a prompt to look closer, never a substitute for looking. Use it to decide where to ask a question, not to decide the answer. A high score justifies a conversation — "walk me through how you approached this" — and a genuine author can almost always talk about their own work in a way no score can. Because the false-positive cost falls entirely on a real student, the responsible default is to treat any single percentage as insufficient grounds for an accusation on its own. Combine it with what you know about the student's prior work, their ability to discuss the piece, and plain common sense. A detector is one input among several, and the weakest one to hang a judgment on.
If you're a writer being screened
Writers increasingly get run through detectors by clients, editors, or platforms, and human writing does get flagged — especially clean, professional, well-structured prose, which is exactly what good writers produce. If your own work scores high, the fix is rarely to make your writing worse. It's to be able to show your process and, where you can, to understand which stylistic traits tend to read as machine-like so you're not blindsided. Knowing that formal, uniform, low-variation writing runs hot on most detectors turns a mysterious flag into an explainable one.
The honest ceiling on what any percentage can prove
Everything above points to one conclusion, and it's worth stating without hedging: no detector percentage, from any tool, at any value, is proof. Not 60%, not 95%, not a flat 100%. A percentage is one tool's probabilistic guess about the origin of a piece of text, rendered in units that make it look far more authoritative than it is. That's genuinely useful as a signal — a place to look, a question to ask, a reason to pay attention. It is not evidence in the sense that word deserves.
The reasons stack up. The number means different things across tools. The same text scores differently depending on where you paste it. The threshold that turns a raw score into a "flag" is a vendor's tunable choice, not a fact. False positives are real, they fall on innocent people, and they're the expensive kind of error. The precision of the display vastly outruns the accuracy underneath it. And the extremes of the scale — 0% and 100% — are the least trustworthy points of all. If you want to see how these limits show up in measured performance rather than in principle, what the accuracy data actually shows puts numbers to the pattern.
None of this means detectors are worthless or that you should ignore them. It means you should read the number for what it is: a compressed, styled, uncertain estimate that a person still has to interpret. The percentage doesn't make the decision. It just tells you a decision might be worth making — carefully, with more than a number in hand, and with the humility to remember that the confident-looking digits on screen were never as certain as they looked. Treat the score as the opening of an inquiry, and it can genuinely help. Treat it as the closing of one, and it will eventually help you make a mistake you can't take back.