What “Is This AI-Generated Text?” Really Means
If you are asking whether a particular passage was written by artificial intelligence, the honest answer is usually probabilistic rather than certain. Text alone rarely contains a reliable, universal marker that proves its author. A human can write in a style that resembles an AI system, and an AI-generated draft can be substantially revised until it resembles ordinary human writing.
No detector, language model, or expert can establish authorship with certainty from an arbitrary passage alone. A credible assessment should combine the text with contextual evidence such as drafts, revision history, document metadata, source notes, writing habits, and—where available—authenticated provenance information.
If no passage has been provided, it is not possible to determine whether “this text” is AI-generated. If a passage is available, it can be examined for indicators associated with generated writing, but those indicators should be described as evidence or likelihood—not as proof.
Why AI-generated text can be difficult to identify
Modern language models generate text by predicting plausible continuations based on patterns learned from large collections of language. Their output often has recognizable tendencies, especially when the system is asked to produce a generic explanation quickly. However, those tendencies are not exclusive to AI.
A human writer may produce text that is:
- Grammatically polished and consistently structured
- Formal, neutral, or impersonal in tone
- Organized around predictable headings
- Repetitive in its use of transitions
- Broad in its claims but light on concrete experience
- Carefully balanced, with few strong opinions
- Free of obvious spelling or punctuation errors
These qualities may reflect editing, professional training, a formal assignment, a non-native writing style, or the use of grammar and translation software. They do not demonstrate that a language model wrote the text.
The reverse is also true. AI-generated text may contain personal details, unusual phrasing, errors, slang, citations, or an apparently distinctive voice. A person may ask an AI system to imitate a particular style, introduce mistakes, or rewrite the output extensively. As a result, surface clues become less useful after editing or paraphrasing.
The important distinction is between authorship evidence and style evidence. Style evidence can suggest that a passage deserves closer examination. It generally cannot identify the author by itself.
Common signs people associate with AI writing
When someone asks, “Is this text AI?” they are often noticing a combination of stylistic features rather than a single telltale sign. The following patterns can be relevant, but each also occurs in human writing.
Highly predictable organization
Generated answers frequently follow a clear template:
- A definition or direct answer
- A series of benefits, causes, or steps
- A balanced discussion of limitations
- A concluding restatement
This structure is useful and often deliberately requested by users. It is not evidence of AI authorship on its own. Textbooks, help articles, workplace reports, and edited essays commonly use the same organization.
Generic specificity
AI writing may sound specific while avoiding verifiable details. It can describe “several important factors” or “a wide range of benefits” without naming people, dates, measurements, examples, or sources. This can create a polished but somewhat interchangeable quality.
A human writer working from limited knowledge can produce the same effect. To evaluate the issue, look for whether the passage makes claims that should be supported and whether its supporting details are accurate and independently checkable.
Repeated transitions and balanced phrasing
Phrases such as “however,” “in addition,” “it is important to note,” and “on the other hand” are common in generated text. So are constructions that present two sides in nearly symmetrical language.
Repeated use of these patterns may be a useful stylistic observation, particularly when it differs sharply from the writer’s earlier work. It remains weak evidence because these are ordinary features of formal prose.
Smoothness without a clear individual voice
Some generated passages have an even rhythm and a consistently moderate tone. They may avoid personal commitments, unusual metaphors, local references, and abrupt changes in emphasis. The result can feel polished but impersonal.
This characteristic is especially unreliable for short passages. A short school response, business memo, or edited professional document may naturally have little individual voice. Conversely, a longer AI response can be made highly personal through prompting and revision.
Unnatural repetition
A language model may restate the same point in several forms, repeat a conclusion, or use several headings that overlap in meaning. It may also explain familiar concepts at length while omitting a key detail that an informed human would normally include.
Repetition is more informative when it appears alongside factual inconsistency, fabricated references, or a mismatch with the assignment. Even then, it indicates a need for review rather than a definite origin.
Incorrect or invented references
A generated passage may cite a real author or publication but attribute the wrong claim to it. In some cases, it may produce a citation, quotation, case, study, or web address that does not exist. These errors are often called hallucinations, meaning unsupported content presented as if it were factual.
A fabricated reference can show that the text was produced or assisted by a system that generated unsupported material, but it does not establish how much of the passage was written by a person. A human could also have copied or invented a citation. Every important reference should be checked directly.
Why AI detectors are not definitive
AI detectors generally estimate whether a passage resembles text produced by language models. Different tools use different signals, such as word-probability patterns, sentence variation, classifier models, or comparisons with known samples. Their output is not a direct observation of authorship.
Two concepts are particularly important:
- A false positive occurs when human-written text is classified as AI-generated.
- A false negative occurs when AI-generated text is classified as human-written.
Both can occur because detector performance depends on the language, length, subject, model family, prompting method, editing, paraphrasing, and the particular detector used. A passage may receive different results from different tools, and a detector may change behavior as new models and writing practices appear.
Short passages are especially difficult to assess because they contain fewer stylistic signals. Text written by children, non-native English speakers, or people using a highly formal style may also be disproportionately vulnerable to false positives. Conversely, lightly edited or paraphrased generated text may evade detection.
OpenAI discontinued its own AI Classifier in July 2023, stating that it had a low rate of accuracy. That decision illustrates a broader limitation: even a detector developed by an organization with expertise in language models should not be treated as an authorship test.
A detector score should therefore be interpreted as a screening signal, not a verdict. In an educational, employment, legal, or disciplinary setting, acting on a score alone can produce an unjust accusation.
A more reliable way to investigate a passage
A careful review should begin with the writing process rather than the prose alone.
1. Preserve the original material
Keep the document in its original form, including the file, timestamps, revision history, and any available comments. Copying text into a new document or online detector can remove useful context and metadata.
Metadata is not conclusive. It can be changed, stripped during export, or altered by collaboration software. It is supporting evidence, not proof.
2. Compare the passage with known writing
If the question concerns a particular author, compare the passage with several verified samples written under similar conditions. Examine:
- Vocabulary and sentence length
- Punctuation and spelling habits
- Typical level of detail
- Use of personal examples
- Organization and paragraph length
- Recurring errors or preferred expressions
- Knowledge and terminology appropriate to the author
A comparison is most useful when the samples are comparable. A casual message should not be treated as a direct baseline for a formal essay, and a heavily edited document should not be compared with an unedited draft.
3. Check facts, sources, and quotations
Verify names, dates, technical claims, quotations, calculations, and references. Unsupported or fabricated details are important quality problems regardless of who wrote the text.
Factual accuracy cannot prove human authorship. Humans make mistakes, and AI systems can produce accurate text. Fact-checking helps identify whether the passage is trustworthy, not whether it came from a particular source.
4. Examine the writing process
Process evidence is often stronger than stylistic intuition. Relevant evidence may include:
- Earlier drafts
- A sequence of revisions
- Notes or outlines
- Research materials
- Document version history
- Comments explaining changes
- The author’s ability to explain the argument and choices
- Evidence of how sources were selected and used
The absence of drafts is not proof that AI was used. Many people write directly in a final document, and some platforms do not preserve a usable history.
5. Ask about assistance, not only authorship
“AI-generated” can describe several different situations:
- A system produced the entire first draft.
- A person used AI to brainstorm or outline.
- A tool rewrote sentences or corrected grammar.
- A person translated or paraphrased machine-generated material.
- AI-generated content was copied without meaningful revision.
- A human wrote the passage independently.
These cases may have different rules in a classroom, workplace, publication, or contract. A fair policy should define what kinds of assistance are permitted and how they should be disclosed rather than relying on an undefined idea of “AI text.”
Provenance and watermarking
The strongest evidence of origin comes from a trustworthy record created during production. Provenance means information about where a digital asset came from and what happened to it over time.
Some systems and publishing workflows can attach signed metadata or other provenance signals. The Coalition for Content Provenance and Authenticity (C2PA), for example, defines an open standard for recording an asset’s origin and editing history through Content Credentials. Such credentials can help show that a participating tool or workflow was involved.
Provenance has important limits:
- It may not be present in every file or platform.
- Exporting, copying, screenshots, or reformatting can remove it.
- It can usually establish that a particular tool handled content, not necessarily that the tool wrote every word.
- A missing credential does not prove that AI was not used.
- A credential may show origin for an earlier version rather than the exact text being examined.
Text watermarking is a related idea in which a generation system subtly favors certain word or token patterns. A detector can then test for the pattern. Watermarks may be weakened or removed by translation, paraphrasing, editing, truncation, or formatting changes, and they require cooperation from the generating system. They are not a universal solution to authorship disputes.
How to interpret a detector result
If you use an AI detector for an informal first review, treat the output cautiously:
- Use more than one method, but do not assume agreement makes the result conclusive.
- Provide enough text for the tool to analyze, while avoiding confidential or sensitive material.
- Record the tool, date, settings, and exact passage.
- Check whether the tool reports a probability, classification, or merely a similarity score.
- Do not interpret a high score as proof or a low score as proof of human authorship.
- Seek process evidence and ask the writer for an explanation before reaching a conclusion.
Do not upload private academic work, legal documents, unpublished manuscripts, personal data, or confidential business material to an unfamiliar service without understanding its retention and usage policies.
What to do when the stakes are high
For low-stakes editing or curiosity, stylistic analysis can help identify passages that might benefit from fact-checking or revision. For high-stakes decisions—such as grading, hiring, publication, professional discipline, immigration, or legal proceedings—general AI-detection advice is not enough.
A decision should rely on a transparent policy, multiple forms of evidence, an opportunity for the writer to respond, and review by the responsible institution or a qualified professional where appropriate. The central question should also be clear: is the concern unauthorized assistance, plagiarism, inaccurate content, undisclosed use, or a breach of confidentiality? Those are different issues.
The most accurate answer to “is it AI text?” is therefore usually conditional: the passage may contain signals associated with AI-generated writing, but text-only analysis cannot reliably prove who wrote it. Establishing origin requires corroborating evidence from the writing process, verified sources, document history, or authenticated provenance—not confidence based on polished wording alone.
Determining Whether Text Is AI-Generated
Evaluating whether a piece of writing is AI-generated text requires analyzing a combination of statistical patterns, stylistic artifacts, and factual coherence. While no single test or automated detector provides 100% certainty, large language models (LLMs) leave distinct linguistic footprints. These models generate text by predicting the mathematically most probable next word (or token) given preceding context, which often results in writing that is unusually uniform, structurally predictable, and stylistically neutral compared to human prose.
To determine if text is synthetic, analysts, educators, and editors generally examine three distinct layers:
- Statistical metrics: Quantifiable measures such as perplexity (how predictable the word choices are) and burstiness (the variation in sentence length and structure).
- Linguistic and structural cues: Overused transition words, formulaic paragraph organization, excessive hedging, and generic summaries.
- Factual and contextual consistency: Superficial depth, hallucinated citations or dates, and an absence of personal, idiosyncratic perspective.
The Statistical Foundations of Synthetic Text
Large language models generate text based on probability distributions over a fixed vocabulary. This probabilistic generation creates measurable statistical signatures that contrast sharply with natural human communication.
Perplexity and Token Predictability
Perplexity measures how well a probability model predicts a sample of text. In the context of AI text detection, it quantifies how "surprised" a language model is by the sequence of words.
If a model assigns high probability to every consecutive token in a passage, the perplexity score is low. Mathematically, for a token sequence , perplexity is defined as the exponentiated average negative log-likelihood:
- Low Perplexity: Indicates that the text chooses the most conventional, standard, and mathematically probable next words. Most default LLM outputs exhibit low perplexity.
- High Perplexity: Indicates unconventional word pairings, novel metaphors, slang, or non-standard syntax common in creative or highly specialized human writing.
Burstiness: Sentence and Structure Variation
Burstiness refers to the variance in sentence length, rhythm, and complexity across a text.
Humans naturally write with high burstiness. An author might write a short, punchy sentence. Then, they might follow it with a long, meandering sentence containing multiple subordinate clauses, parenthetical thoughts, and rhythmic shifts.
In contrast, standard LLM outputs tend to maintain a steady, moderate cadence. Sentences are typically between 15 and 25 words long, use parallel grammatical structures, and follow a predictable Subject-Verb-Object rhythm throughout the entire piece.
| Feature | Human-Authored Text | AI-Generated Text (Standard) |
|---|---|---|
| Perplexity | Highly variable; ranges from predictable to idiosyncratic. | Consistently low to moderate; avoids rare or jarring token combinations. |
| Burstiness | High variance; dramatic mix of short and long sentences. | Low variance; uniform sentence lengths and rhythmic pacing. |
| Vocabulary Distribution | Uneven; includes regionalisms, non-standard idioms, and domain shorthand. | Broad but standard; heavily favors high-frequency, neutral terminology. |
| Punctuation Diversity | Irregular use of dashes, semicolons, exclamation marks, or informal ellipses. | Strict adherence to standard comma rules, bullet points, and colons. |
Common Linguistic Markers and Stylistic Habits
Beyond mathematical metrics, raw AI text frequently exhibits recurring rhetorical patterns and stylistic "tics." These arise from the instruction-tuning and Reinforcement Learning from Human Feedback (RLHF) processes used to align modern models.
[Typical AI Structural Flow]
┌─────────────────────────────────────────────────────────────┐
│ 1. Broad Definitional Opening / Restatement of the Prompt │
├─────────────────────────────────────────────────────────────┤
│ 2. Symmetrical Body Paragraphs (often introduced by │
│ "Furthermore," "Moreover," "Additionally") │
├─────────────────────────────────────────────────────────────┤
│ 3. Balanced Perspective ("While X has benefits, Y presents")│
├─────────────────────────────────────────────────────────────┤
│ 4. "Crucial/Tapestry/Beacon" Metaphorical Tropes │
├─────────────────────────────────────────────────────────────┤
│ 5. Concluding Paragraph beginning with "In conclusion," or │
│ "Ultimately, [topic] serves as a reminder that..." │
└─────────────────────────────────────────────────────────────┘Overused Vocabulary and Tropes
RLHF training optimizes models to produce helpful, polite, and comprehensive answers. This tends to concentrate probability mass around specific vocabulary:
- Adjectives and Verbs: Words such as delve, testament, tapestry, beacon, crucial, pivotal, multifaceted, foster, navigate, underscore, and showcase.
- Formulaic Transitions: Heavy reliance on furthermore, moreover, in summary, on the other hand, it is important to remember, and as an AI.
- Metaphorical Cliches: Overuse of phrases like "a double-edged sword," "a testament to human ingenuity," or "navigating the complex landscape of."
Symmetrical and Parallel Organization
LLMs frequently organize answers into rigid, predictable structures:
- Listicles with Bold Prefixes: An almost compulsive tendency to break explanations into bulleted lists where every item begins with a bolded term followed by a colon and a two-sentence elaboration.
- The "Bookend" Structure: An introductory paragraph that restates the prompt, three body points of near-identical length, and a concluding paragraph summarizing the previous points without adding new insight.
- Excessive Neutrality and Hedging: Models are trained to avoid taking controversial stances unless specifically prompted. This results in constant qualifiers: "While some argue X, others contend Y, suggesting that a balanced approach is essential."
Superficial Depth Without Personal Context
A human subject-matter expert usually references concrete edge cases, practical failures, unwritten industry norms, or personal experiences. AI text often provides wide, accurate conceptual summaries that remain purely theoretical. It describes what a concept is in textbook terms, but rarely demonstrates how that concept breaks down in messy, real-world execution unless explicitly instructed to do so.
How Automated AI Detection Tools Work
Automated detection platforms—such as Turnitin, GPTZero, CopyLeaks, and Winston AI—use several technical strategies to assign a probability score to a piece of text.
1. Classifier Models (LLM-Based Discriminators)
Most advanced detectors are themselves language models (often RoBERTa or custom transformer models) trained on paired datasets containing millions of human-written and AI-generated texts. The classifier learns high-dimensional feature representations that distinguish human syntactic structures from machine-generated patterns.
2. Zero-Shot Perplexity Mapping
Rather than using a trained classifier, some tools pass the candidate text through a reference model (such as GPT-2 or a small open-source LLM) to calculate the exact probability of every token. If the reference model predicts the author's next word with high accuracy across 90% of the text, the detector flags the passage as synthetic.
3. Watermarking Schemes
Some AI developers embed cryptographic statistical watermarks directly into the text generation process.
A watermarking algorithm partitions the model's vocabulary into "green" and "red" token lists based on the hash of preceding tokens. During generation, the model artificially biases sampling toward the green list. While invisible to a human reader, a detection algorithm that knows the hashing key can analyze a passage and calculate whether the proportion of green tokens exceeds what could occur by random chance.
The Problem of Reliability: False Positives and Negatives
Automated AI detectors are not definitive forensic proof. They output a statistical probability, not a deterministic verdict. Relying solely on software scores creates serious risks in educational, academic, and professional environments.
┌───────────────────────────────┐
│ Automated Detector Score │
└──────────────┬────────────────┘
│
┌───────────────┴───────────────┐
▼ ▼
┌────────────────────┐ ┌────────────────────┐
│ False Positive │ │ False Negative │
└──────────┬─────────┘ └──────────┴─────────┘
│ │
• Non-native English writing • Heavy prompting / persona
• Technical & legal writing • Paraphrasing & human edits
• Standardized formulas • Temperature adjustmentsCauses of False Positives
A false positive occurs when purely human-written text is flagged as AI-generated. This happens most frequently in:
- Non-Native English Writing: Studies have shown that essays by non-native English speakers are disproportionately misclassified as AI-generated. Non-native writers often use simpler sentence structures, standard transition words, and a more constrained vocabulary, which lowers the perplexity score.
- Formulaic and Technical Prose: Legal contracts, medical case reports, standard lab protocols, and corporate compliance documents follow strict, conventional patterns. Because these formats prioritize standardization over stylistic variety, detectors frequently label them as synthetic.
- Short Text Samples: Passages with fewer than 250–300 words do not provide enough statistical data for reliable perplexity and burstiness analysis, drastically increasing the error rate.
Causes of False Negatives
A false negative occurs when AI text bypasses detection. Generation can be disguised through several methods:
- Prompt Engineering: Instructing an LLM to "write with high burstiness, mix sentence lengths, avoid cliches like 'delve' or 'tapestry,' and include occasional sentence fragments" significantly alters statistical distributions.
- Adjusting Sampling Parameters: Raising the sampling temperature or top-p value forces the model to select less probable tokens, increasing perplexity.
- Paraphrasing Tools and Hybrid Editing: Running AI output through secondary spinning tools or having a human lightly edit every third sentence disrupts the continuous token sequences that classifiers search for.
A Practical Framework for Evaluating Text
When evaluating whether an essay, article, or report is machine-generated, combine automated tools with human editorial review.
[Review Workflow]
1. Baseline Screening ──► Run 2+ reliable detectors to check statistical bounds.
2. Anomaly Checking ──► Inspect references, citations, dates, and quotes.
3. Stylistic Inspection ──► Look for structural symmetry, RLHF cliches, and low burstiness.
4. Human Verification ──► Compare against the author's known writing baseline.Step 1: Check Facts and Citations
The most definitive evidence of synthetic text is not stylistic—it is factual:
- Hallucinated References: Search for cited book titles, academic papers, DOIs, or legal cases. AI models frequently synthesize plausible-sounding academic titles paired with real author names that never actually published such works.
- Anachronisms and Logical Drift: Check for claims that conflate distinct historical events or present contradictory points within adjacent paragraphs.
Step 2: Compare Against Known Author Baselines
Stylistic analysis is most reliable when compared against verified samples of the author's previous work. Look for sudden, inexplicable shifts in:
- Vocabulary range and syntactic complexity.
- Grammatical precision and punctuation habits.
- Domain-specific terminology usage versus generic explanatory phrasing.
Step 3: Conduct an In-Person or Contextual Check
In institutional or academic settings where authorship integrity is critical, software flags should only serve as an investigative trigger, not final evidence. The most reliable follow-up is a brief conversation asking the author to:
- Explain their research methodology and draft evolution.
- Define specialized vocabulary used within the text.
- Provide draft revision history, outlines, or research notes.
The honest answer: no tool can tell you for certain
There is no reliable way to look at a piece of text — by eye or with software — and prove that it was or was not produced by a language model. If you are asking "is this AI generated text?" about a specific paragraph, essay, email, or article, the best available answer is a probability estimate, and that estimate is often much weaker than detector interfaces make it appear. Detectors output confident-looking percentages, but those numbers describe statistical similarity to model-like writing, not evidence of authorship.
This is not a temporary gap that better tools will soon close. Modern language models are trained specifically to produce text that resembles competent human writing, so the distributional gap detectors exploit keeps narrowing. OpenAI withdrew its own AI Text Classifier in July 2023, citing a low rate of accuracy — a telling signal that the organization with the most direct knowledge of how its models write could not build a dependable detector for them.
That said, "you can't be certain" is not the same as "you know nothing." There are several distinct ways to reason about whether a text is AI-generated, and they differ enormously in reliability.
Four different ways to ask the question
It helps to separate the approaches, because people often confuse a weak method with a strong one.
1. Statistical detection (weak to moderate). Classifiers and zero-shot methods estimate how "predictable" text looks to a language model. The classic signals are perplexity (how surprised a model is by the next word) and burstiness (how much sentence length and complexity vary). AI text tends to sit in a smooth, high-probability region of the model's distribution; human writing tends to be lumpier. Research methods such as DetectGPT and its faster successors refine this by examining the curvature of the probability landscape — machine text usually sits near a local maximum, so small perturbations lower its likelihood more sharply than they do for human text. Commercial detectors typically train a supervised classifier on large corpora of human and machine text instead. All of these are statistical inferences about a population, applied to an individual document.
2. Watermarking (strong, but only when present). A watermark is deliberately embedded at generation time by biasing the model's token sampling in a pattern that a matching detector can recognize. Google DeepMind's SynthID-Text, described in a 2024 Nature paper and released as an open reference implementation, is the most prominent example. Watermark detection is far more trustworthy than post-hoc classification — but only for text generated by a system that applied that specific watermark, and only if the text hasn't been heavily rewritten or translated. Most text on the internet carries no watermark at all, so a negative result means almost nothing.
3. Provenance and metadata (strong when available). Document version history, editing timelines, draft folders, C2PA-style content credentials, and platform-side generation logs are usually far more informative than any text analysis. A Google Doc that materialized in one paste at 2:14 a.m. tells you something a perplexity score cannot.
4. Contextual and human judgment (variable). Domain knowledge, comparison to a person's prior writing, factual errors, and fabricated citations often reveal more than a detector does — but they also produce confident mistakes.
How to read a detector score without fooling yourself
Suppose a tool reports "92% likely AI." That number is not the probability that the text is AI-generated. It is, at best, a classifier confidence calibrated on some benchmark dataset that may look nothing like your document.
The decisive missing ingredient is the base rate: how common AI text actually is in the pool you're sampling from. Even a good detector produces misleading results when genuine AI text is rare. Take a detector with a 95% true-positive rate and a 4% false-positive rate — figures roughly in line with what vendors like Turnitin have publicly discussed at the sentence level — applied to a population where 5% of documents are AI-written:
Barely better than a coin flip, despite headline accuracy that sounds excellent. Flip the base rate to 50% and the same detector becomes far more convincing. This is why the same tool can be reasonable for triaging a bulk content farm and completely inappropriate for accusing one student.
Other practical cautions when interpreting a score:
- Length matters enormously. Under roughly 150–300 words, statistical signals are noisy and detectors become close to unusable. Single sentences are essentially undetectable.
- Genre matters. Formulaic writing — lab reports, legal boilerplate, technical documentation, SEO content, standardized-format essays — is naturally low-perplexity and gets flagged more often, regardless of who wrote it.
- Detectors disagree. Running text through several tools frequently produces contradictory verdicts. Agreement across tools is weak corroboration, not proof, because many share similar underlying assumptions.
- Scores are not stable over time. Vendors retrain models; the same text can score differently months apart.
Why false positives land on specific people
The fairness problem is well documented. A 2023 study published in Patterns found that GPT detectors frequently misclassified writing by non-native English speakers as machine-generated, because simpler vocabulary and more conventional sentence construction mimic the low-perplexity profile detectors key on. Stanford HAI and numerous university teaching centers have echoed the finding.
The people most likely to be wrongly flagged include:
- Non-native English writers and multilingual authors
- Writers using grammar assistants, translation tools, or predictive text
- Neurodivergent writers with highly regular sentence structures
- Anyone writing in a rigid template or house style
- Students taught to write in a formulaic five-paragraph structure
Meanwhile, false negatives cluster among the sophisticated: lightly paraphrased output, text passed through "humanizer" tools, output generated with unusual sampling settings or system prompts, and hybrid drafts where a person rewrote machine text. Research on paraphrase attacks consistently shows detection accuracy degrading substantially, and watermarks weaken under aggressive rewriting. The result is an uncomfortable asymmetry: detectors are hardest on careless honest writers and easiest on careful dishonest ones.
Textual clues people rely on, and what they're actually worth
Human intuition about AI text has some signal, but it decays quickly as models improve and as writers absorb model style.
| Clue | Reliability | Why it misleads |
|---|---|---|
| Fabricated citations, fake DOIs, invented case law | High when verified | Requires checking; humans also cite sloppily |
| Confident factual errors about verifiable specifics | Moderate–high | Humans make errors too, but hallucinations have a distinctive plausible-but-wrong texture |
| Uniform paragraph length, tidy triads, symmetrical structure | Moderate | Also characteristic of trained professional and academic writing |
| "It's not just X, it's Y", "delve", "tapestry", "in today's fast-paced world" | Weak and fading | These are stylistic fashions humans now imitate; models also shift phrasing across versions |
| Em dashes, bulleted summaries, hedged conclusions | Very weak | Long-standing features of edited prose |
| Absence of typos, absence of personal detail | Weak | Describes copyediting, not authorship |
| Missing lived specificity where it would be expected | Moderate | Useful in memoir, reviews, first-person reporting |
The strongest human-side check is rarely stylistic at all: it is verification. Do the sources exist? Do the numbers reconcile? Does the writer know the material when asked an unscripted follow-up question? Fabrication is checkable; style is not.
If you have been accused of using AI
General guidance, not legal advice — institutional policies vary widely and some outcomes carry real consequences worth professional review.
- Ask what the evidence is. A detector percentage alone is a weak basis for an allegation, and many institutions' own policies say scores should not be used as sole evidence.
- Produce process evidence. Document version history, draft files, browser or editor timestamps, notes, outlines, library records, and search history are far more persuasive than arguing about the score.
- Offer to discuss the content. A short oral conversation about sources, argument choices, and revisions usually resolves the question quickly in either direction.
- Ask about the tool's stated limitations. Vendors publish false-positive caveats; those documents are legitimate to cite.
- Check whether assistive tools count. Many policies distinguish spellcheck and grammar assistance from generation. Clarify which you used.
If you are the one evaluating text
For teachers, editors, hiring managers, and platform moderators, the practical stance that holds up is to treat detection as a triage signal, never a verdict.
- Use scores to decide where to look more closely, not what to conclude.
- Pair any flag with process evidence or a conversation before acting.
- Set the decision threshold according to consequence: a high-stakes accusation demands far more than a moderate score.
- Redesign the task where possible. Assignments and briefs that require specific local knowledge, personal reflection, in-class or live work, iterative drafts, or engagement with a very recent source are more resilient than any detector.
- Be explicit in policy about what disclosure you require, since "AI-assisted" covers everything from grammar fixes to wholesale generation.
The hybrid problem and where this is heading
The framing "is this text AI or human?" is increasingly the wrong question, because most professional writing now sits on a spectrum: human-outlined and machine-drafted, machine-drafted and human-rewritten, human-written and machine-polished. Detectors are built as binary classifiers and have no coherent answer for a paragraph that a person wrote and a model tightened. Sentence-level highlighting gives an illusion of precision that the underlying statistics do not support.
The plausible long-run direction is provenance rather than detection — signed metadata, content credentials, platform-side generation records, and watermarking applied at the source. Regulation is pushing that way: Article 50 of the EU AI Act imposes transparency obligations on providers of systems generating synthetic content, including marking outputs in a machine-readable format, with those provisions applying from 2 August 2026 and accompanying guidance and a code of practice under development. Such regimes will make some AI text identifiable by declaration. They will not make unmarked or stripped text detectable, and they cannot retroactively label the enormous volume of machine text already in circulation.
So the durable answer to "is this text AI generated?" is that you can often gather evidence, sometimes gather strong evidence when a watermark or provenance trail exists, and almost never obtain proof from the words alone. Treat any tool that claims otherwise — in either direction — as overstating what the underlying method can support.