"not X, but Y" pivot
HIGHWEIGHT 1.0PATTERNThe negation pivot is one of the strongest single predictors of machine text. Say the positive thing directly.
Every pattern this tool looks for, printed in full: what it is, why it reads as machine-made, an example sentence, and the rewrite the engine produces. Other humanizers keep the list behind the button. This is the list.
The entries are not a description of the registry. They are the registry, imported into this page. The examples below are scored by the same analyzer that runs in the editor, in your browser, as this page loads — which is also why the highlight sits exactly where the detector fired and not a character to either side.
A tell is a specific, nameable pattern that turns up far more often in generated text than in human drafts. Not a vibe, not a hunch about tone: a construction you can point at, quote, and argue with. "Not X, but Y." Three items where two would do. A transition that connects nothing. Six sentences of almost identical length.
None of these are errors. Every one is grammatical, and most of them appear somewhere in writing you admire. What makes them a signal is rate. A language model produces text by choosing high-probability continuations, and these constructions are high-probability continuations almost everywhere — which is exactly why they accumulate in a way human drafting does not.
That is also the honest link to detection. Commercial detectors do not hold a list like this one and check your paragraph against it. They score predictability: how surprised a language model is by each token you wrote, and how much your sentence lengths vary around their own mean. The patterns below are what low surprise looks like when you read it instead of measuring it. Cut them and the measurable thing usually moves too, because you have replaced probable text with specific text.
Which means a tell is evidence, never proof. Plenty of people write this way naturally, and some of them get flagged for it. Stanford HAI's 2023 study (Liang et al., arXiv 2304.02819) found that 61% of TOEFL essays written by non-native English speakers were falsely flagged as AI across seven detectors. Every one of those essays had a human author. Careful, textbook-correct, evenly paced English is precisely what a predictability classifier scores as machine-made.
So read the list two ways. If you are editing a permitted AI draft, it is a checklist of what to cut. If you have been accused, it is a map of which of your own habits are carrying the risk — and the honest answer is that some of them are just how you write.
Each detector below carries four pieces of machine-readable data and one sentence of plain-language reasoning. All five come out of the registry, not out of this page.
High or medium. High-severity tells are strong enough to matter on their own; medium ones only count when they cluster. The band drives how the finding is ranked when two patterns overlap the same words.
The contribution to weighted tell density, between 0.5 and 1. Density saturates rather than stacking linearly, because the tenth tell in a paragraph tells you much less than the second.
41 detectors are regular expressions over the text. 3 are whole-document functions that need every sentence at once — those are the cadence readings, and they report measured numbers.
Where a mechanical fix exists, the engine offers candidates that keep your inflection. Where none exists, the entry says so. An empty candidate means "cut this", and the applier repairs the spacing and capitalisation.
These are shapes rather than words. A language model builds a sentence by picking the most probable continuation, and the most probable continuation is very often a symmetry: a negation followed by its correction, a list that resolves in three. The constructions are all perfectly good English. What marks them is the rate. A human writer reaches for the rule of three when the argument has three parts; generated prose reaches for it because the shape itself scores well.
The negation pivot is one of the strongest single predictors of machine text. Say the positive thing directly.
Same negation reflex in clause form. Readers only need the claim you actually mean.
Language models reach for the rule of three constantly. Two specifics beat three abstractions.
Three gerunds in parallel is textbook generated cadence. Pick the one that actually happened.
A correlative pair that adds emphasis without adding information. Drop the scaffolding.
A conclusion that announces its own significance instead of showing it.
Readers can see it is the last paragraph. The label is pure filler.
Another negation pivot wearing a different hat.
Two words of throat-clearing in front of a plain copula.
Asking a question you immediately answer is a generated transition, not a real turn in the argument.
No mechanical rewrite. This tell is flagged and explained, and the decision about what to do stays with the writer.
A restatement frame. Say the thing once, in the sentence it belongs to.
First sentences are where generated text is most predictable, because a model with no context yet has nothing to condition on except the register of the prompt. The result is scene-setting that describes no scene, or a claim of importance made before anything has been claimed. If you cut the opener and the paragraph still starts correctly, the opener was ornament.
Scene-setting that describes no scene. Detectors weight it heavily because almost nothing else opens this way.
Four words that mean "in".
Cosmic throat-clearing. Start where your evidence starts.
The fix is a deletion, and deleting alone leaves a fragment here: "the printing press, gatekeepers have set the terms." Some tells need the sentence re-joined by hand, and the editor says so rather than pretending otherwise.
A five-word preposition. "For" or "with" carries the same load.
The single most over-produced clause in generated prose.
Ceremonial register. Name what the thing proves.
A promise of essence that the sentence never keeps.
A connective earns its place by naming a relationship the reader could not infer. Most of the ones below name nothing: they sit at the head of a sentence that already follows from the one before it. Generated text over-produces them because formal-register connectives are cheap probability — they fit almost anywhere, so they are almost always an available next token.
These connect nothing. If the next sentence follows, it follows without an usher.
Formal-register connectives cluster in generated text far more than in human drafts.
A contrast marker in front of a sentence that already contrasts.
Ordinal adverbs are rare in natural prose and common in generated outlines.
If it were not worth noting you would not have written it.
Word-level habits, and the closest thing here to a fingerprint. Some of these are ordinary words used at an extraordinary rate: "delve" is normal English that appears in post-2023 generated text far more often than in human drafts. Each entry carries inflection-aware replacements, so the rewrite keeps your tense rather than dropping a dictionary form into the middle of a sentence.
Post-2023 generated text uses this verb at many times the human rate. It is close to a fingerprint.
Consultant-register verb that models default to. "Use" says the same thing.
Almost always replaceable with a concrete verb — build, grow, encourage.
A longer word for "use" with no extra meaning attached.
A rating word standing in for a description. Say what makes it hold up.
Marketing adjective. If nothing broke, say nothing broke.
Self-praise adjective that models attach to any process noun.
A stock collocation that survives from prompt to output almost unchanged.
Quantity words that avoid naming a quantity.
Tapestry, synergy, paradigm: nouns that name a vibe instead of a thing.
Borrowed ecology words standing in for the actual noun. Sometimes right — usually a placeholder.
A promise with no verb of its own. Say what the thing can now do.
A verdict borrowed from press releases. Name the change instead.
Rating adjectives inflate without informing. Cut, or replace with the stake itself.
Hedging is how a model avoids committing to a claim it cannot check. Stacked hedges are the giveaway: two qualifiers in a row cancel each other, and the sentence ends up asserting nothing while sounding careful. Cutting the hedge usually improves the writing whatever wrote it — which is a useful property for a tell to have.
Two hedges in a row cancel each other and read as risk-averse machine output.
Six words of distance between you and your own claim.
The reader already knows the sentence is yours. Stating the claim is more confident than announcing confidence in it.
A statement with the information removed.
The three detectors here read the whole document rather than a phrase, and they are the ones closest to what statistical detectors actually measure. Rhythm cannot be seen inside a single sentence. It shows up in the distribution: how much sentence length varies, how often consecutive sentences start the same way, how thick the punctuation gets. These produce findings with numbers in them, measured from your own text.
Three or more sentences in a row opening on the same word is parallel structure the reader can hear.
3 sentences in a row open with "the". Vary the entry point: start one on the object, or fold two sentences together.
No mechanical rewrite. This tell is flagged and explained, and the decision about what to do stays with the writer.
Human drafts vary sentence length sharply. A flat length distribution is what perplexity-based detectors score hardest.
Sentence lengths vary by only 0.14 (coefficient of variation) across 6 sentences. Human drafts usually land above 0.45. Break one sentence in half and let another run long.
No mechanical rewrite. This tell is flagged and explained, and the decision about what to do stays with the writer.
Em dashes at this density are a post-2023 generated-prose signature. Commas and full stops do the same work.
This one marks every dash separately and offers a comma or a full stop at each. There is no single rewrite of the passage, because the right substitute differs dash by dash.
Regex sees surface, not meaning. A pattern cannot tell whether your three-item list is three real things or filler, whether "ecosystem" is the right word for a coral reef, or whether you meant the hedge. It marks the shape and hands you the judgement. That is why every finding is explained rather than auto-applied, and why 3 of the 44 deliberately ship with no mechanical rewrite at all.
Absence of tells is not a clean bill. A paragraph with zero registry hits can still read as machine-made through rhythm alone, which is what the three cadence detectors are for — and even they only measure what they measure. A low reading here is not a prediction about any vendor's classifier.
The list ages. Models change, and the habits change with them. "Delve" became a signal after 2023 because of how a particular generation of models was tuned, not because of anything about the word. Expect entries to be added and retired. Cite this page with a date.
None of this is a route past an AI policy. If your institution forbids generative AI in submitted work, editing the output does not change that, and no promise of undetectability would be honest anyway. This registry is built for two jobs: making a permitted AI draft read like a person wrote it, and showing a falsely flagged writer which sentences are doing the damage.
The editor runs this exact registry over whatever you paste and marks each hit where it occurs, with the same explanation you just read and the rewrite candidates alongside it. The checker runs in the browser tab: no upload to score, there is no account, and it keeps working with the network off.
Paste a paragraph you already suspect and read the marks in order. Most drafts have two or three habits repeating rather than forty different problems, and once you can name them you stop writing them.
Not literally. Commercial detectors are statistical classifiers that score how predictable your text is, token by token. This list is the readable surface of that predictability. Commercial detectors are statistical classifiers: they score how predictable your text is, token by token, against a model of ordinary language. What the list below gives you is the readable surface of that predictability. Phrases that are highly probable continuations are exactly the phrases a perplexity-based detector finds unsurprising, which is why removing them tends to move a score. The registry is ours. The weights inside a vendor classifier are theirs, and unpublished.
No. Every pattern here occurs in human writing, some of it very good. A single "moreover" is a word choice. What separates a generated paragraph from a human one is density and co-occurrence: several tells at once, in a short span, on top of a flat rhythm. That is why each entry carries a weight rather than a verdict, and why the score is a density reading rather than an accusation.
Because the patterns are properties of English generated text and we have not built or checked the equivalent lists for other languages. The two rhythm measurements — sentence-length variation and repeated openers — are close to language-agnostic and will still read a non-English draft. The named tells will not fire.
Yes, and you do not need to ask. Cite it as the Humanize AI tell registry with the date you read it, because it changes as detectors and models change. If you are writing about detection and want the reasoning behind a specific entry, every entry on this page is rendered directly from the registry the analyzer uses, so what you read here is what runs on this page.
Cutting these constructions lowers the reading this tool gives you. Detectors are retrained without notice and disagree with each other on the same paragraph. What we can say is what this page shows: these are specific, nameable habits, most of them make the writing worse, and cutting them usually lowers a detectability reading. That is a mechanism, not a guarantee, and What a third-party detector reports on a given day is its own call, on a model it retrains without notice.
Back to Humanize AI, or read the Detector Observatory.