Can you actually tell if it’s AI?

Almost everyone believes they can spot AI writing. The em dashes. The “it’s worth noting.” The suspiciously balanced conclusion that begins “In conclusion.”

The research says otherwise. Across more than half a dozen controlled studies summarised in a 2025 Journal of Artificial Intelligence Research survey by Fraser, Dawkins and Kiritchenko of the National Research Council Canada, unaided human accuracy at distinguishing AI-generated text from human-written text clusters in a narrow and uncomfortable band: roughly 50% to 61% – that is, from coin-flip to barely better than coin-flip.

Anyone making decisions about content provenance really needs to chew on this number because the intuition “I’ll just read it and know” is the single most common alternative to actually checking. Here is what the evidence looks like.

The headline numbers

StudyTask / populationHuman accuracy
Liu et al. (2024)Scientific abstracts, general annotators≈ Random guessing
Sarvazyan et al. (2023)Multi-domain (English, Spanish) annotators≈ Random guessing baseline
Uchendu et al. (2021)Single-text and paired “pick the AI” taskChance level in both
Li et al. (2024)Three annotators majoring in linguistics“Only slightly better than chance”
Verma et al. (2024)Six undergrad + PhD students already familiar with AI text59%
Liu et al. (2023b)ESL teachers judging student essays61%, rising to 67% after minimal exposure and self-training
Gehrmann et al. (2019)Human judges, unaided → then given a statistical visualisation tool54% → 72%

Two findings in that table are worth pointing out.

First, Sarvazyan et al. (2023) found no significant difference between annotators experienced with AI-generated text and those with no prior exposure at all. Familiarity with ChatGPT did not translate into detection skill. The confident reader and the naive reader performed about equally.

Second, the trained linguists in Li et al. (2024) performed “only slightly better than chance.” If professional attention to syntax, register and lexical choice buys you a few percentage points, then the folk heuristics most of us use are buying us approximately nothing.

The survey’s authors are blunt about the implication:

“The capabilities of these models to generate realistic, fluent text has exceeded our human ability to detect it as computer-generated.”

Why “I can always tell” feels so true

The feeling isn’t a delusion – it’s a sampling error. You do correctly identify AI text sometimes. What you never see is the denominator.

Every piece of competently generated, lightly edited AI text you read and accept as human passes through you invisibly. You only ever get feedback on the clumsy cases: the unedited first draft, the five-paragraph listicle, the LinkedIn post that opens with a one-word sentence. Your hit rate on obvious AI text may genuinely be high. Your hit rate on the overall population of AI text is what the studies above measure, and it is close to chance.

There’s a second bias baked in. Liu et al. (2024) found that annotators judging scientific abstracts had a systematic tendency to assume everything was human-written. Absent a strong signal, people default to “human.” That default is comfortable, and in an environment where a large share of online text now involves a model somewhere in the pipeline, it is increasingly wrong.

The “tells” that people report –

Human judges aren’t operating randomly. They’re using real cues. Cui et al. (2023) asked participants what they were relying on, and the answers will look familiar:

  • The tone is “overly formal.”
  • Statements are too objective and avoid subjective opinion.
  • High frequency of stock phrases such as “it’s worth noting” or “please note.”
  • Enumerated lists.
  • A formal concluding sentence that summarises what was just said.

Corpus studies partly corroborate this. Nguyen-Son et al. (2017) found humans are more likely than machines to use clichés, idioms, archaic language, contractions like wanna and gonna, and pronunciation spellings like goin’. Muñoz-Ortiz et al. (2023), comparing New York Times articles against Llama-generated news, found the AI text had a more restricted vocabulary, used fewer adjectives, more symbols and numbers, fewer words associated with negative emotions, and a higher propensity for male pronouns.

So the cues are real. The problem is that they are weak, domain-specific and trivially removable.

They assume human writing is the more sophisticated writing. As the survey notes, most stylistic heuristics implicitly assume human text has richer vocabulary, more complex syntax and better coherence. That holds for professional journalism. It does not hold for casual writing, text messages, or social media, where AI text may instead stand out for being “too correct.” Your heuristic has to invert depending on genre, and most readers don’t invert it.

They’re a prompt away from disappearing. “Overly formal” and “lists everything” are properties of default output. A user who asks for a specific voice, or who edits, eliminates them.

They punish the wrong people. This is the most serious issue. Liang et al. (2023) found that texts by non-native English speakers are disproportionately flagged as AI. Using seven public detection tools, accuracy was near-perfect on US 8th-grade student essays, but TOEFL essays written by Chinese English learners produced a false positive rate close to 60%. Ardito (2024) explains why: second-language writers tend toward common vocabulary and simple, formulaic syntax – the same profile as a model optimising for unsurprising next words. Any heuristic built on “this reads a bit flat and formal” is, structurally, a heuristic that mistakes non-native fluency for machine output.

What machines measure that eyes can’t

The reason software outperforms human readers isn’t that it’s cleverer. It’s that it has access to a different class of evidence.

A language model generates text by repeatedly selecting from among its highest-probability next tokens. That process leaves statistical residue that is real but not perceptible to a human reader:

  • Perplexity and entropy. AI text tends to be measurably less surprising than human text relative to a model’s probability distribution. A reader can’t compute this; a detector can.
  • Probability curvature. Mitchell et al. (2023) showed AI text sits in negative-curvature regions of the probability function – small perturbations reliably lower its likelihood. This is the basis of DetectGPT and its faster successor, Fast-DetectGPT (Bao et al., 2024).
  • Burstiness. Variability of style, tone and vocabulary across a document, which is higher in human writing. It’s a core feature in GPTZero’s published methodology (Tian, 2023).
  • Intrinsic dimensionality. Tulchinskii et al. (2023) found human-written text has an intrinsic dimension of roughly 9 to 10, while AI text sits near 8 – a separation that held across genres and across GPT-2, GPT-3.5 and OPT-13B. No amount of careful reading recovers that number.

For calibration on the scale of the gap: OpenAI’s early RoBERTa-based classifier detected GPT-2 output at 95% accuracy (Solaiman et al., 2019), against human performance in the same era measured at 54% unaided (Gehrmann et al., 2019).

Four reasons human detection keeps getting harder

1. Bigger models are less detectable – predictably so. Chakraborty et al. (2023) introduced an AI Detectability Index reflecting that newer, larger models have statistical signatures approaching human distributions. Pagnoni et al. (2022) quantified it: detectability follows a power law, meaning accuracy declines linearly as parameter count grows exponentially.

2. Sampling settings matter more than people realise. Nucleus sampling produces the hardest text to detect. Pu et al. (2023) showed that simply moving a nucleus threshold from 0.96 to 0.8 dropped a statistical detector’s recall by 13%; changing top-k from 40 to 160 dropped recall by 56.4%. None of this is visible to a reader.

3. Length is a hard floor. Li et al. (2024) found detectors need roughly 120 words to reach full potential; He et al. (2024) put it nearer 200 words for ChatGPT-turbo and GPT-4; Chakraborty et al. (2023) proved theoretically that ~500 words should suffice even when human and AI distributions are very close. Below those thresholds – a tweet, a product review, a short forum reply – there is often simply not enough signal, for software or for you. Aggregation helps: Stiff and Johansson (2022) raised a detector from 80% to near 100% by concatenating 10 tweets from the same account.

4. Hybrid text is the real world, and it’s the hardest case. Liu et al. (2024) tested detectors on abstracts that were human-written and then polished by ChatGPT. On that category, the evaluated tools – including GPTZero, ZeroGPT and OpenAI’s classifier – performed worse than random guessing. On the MIXSET benchmark, Zhang et al. (2024b) found that no statistical or fine-tuned detector performed well on mixed human-AI text unless mixed examples appeared in training – though when they did, the RADAR detector reached about 88% across all mixcase categories. AI paraphrasing and “humanizing” (adding typos and errors) were the hardest variants.

And for hybrid misinformation, the picture is worse still. Schuster et al. (2020) found AI-assisted fake news with only a few facts altered is very difficult for both humans and machines to catch — consistent with decades of deception research showing the hardest lies to detect are the ones closest to the truth (Mazar et al., 2008). Worse, Zellers et al. (2019) and Spitale et al. (2023) found AI-generated misinformation is often perceived as more credible and trustworthy than the human-written kind.

When humans do get better

Three interventions do measurably help.

  • Side-by-side comparison. Guo et al. (2023) found detection is easier when judges see a pair – one human, one ChatGPT response to the same prompt – and in that setting, experts familiar with ChatGPT substantially outperformed amateurs. (Note the contrary result: Uchendu et al. (2021) found chance-level performance even in paired “pick the AI” tasks, so this effect is not robust across setups.)
  • Small amounts of deliberate training. The ESL teachers in Liu et al. (2023b) went from 61% to 67% with minimal exposure and self-training. Real, but a 6-point gain from a baseline that is already close to guessing.
  • Giving humans the statistical signal. This is the standout finding. When Gehrmann et al. (2019) gave judges a graphical interface overlaying token-probability information on the text, accuracy jumped from 54% to 72% – an 18-point gain, larger than anything achieved by expertise, training or task design alone.

That last result is the practical thesis of this whole literature. The winning configuration isn’t human or machine. It’s a human reading model-derived evidence they could never have computed themselves.

What to actually do with this

Stop treating your gut as evidence. A 59% instinct is not a basis for an academic integrity case, a contributor ban, or a hiring rejection. If you are going to act on a suspicion, the suspicion needs to be corroborated by something that measures what you cannot see.

But don’t treat any single score as proof either. PlagPointer has over 99% AI detection accuracy and a false-positive rate of just 0.03% but we’d be doing you a disservice to pretend the tools are infallible. In the survey, no off-the-shelf detector tested was a clear winner across evaluations; GPTZero struggled to detect essays from models other than the GPT-3.5 it was calibrated on; DetectGPT did not transfer cleanly between ChatGPT and GPT-3; LongFormer detector lost over 20% accuracy out-of-distribution; and paraphrasing attacks using DIPPER (Krishna et al., 2023) drove true positive rates at a 1% false positive rate down to under 5% for statistical detectors and 13% for a fine-tuned classifier. Weber-Wulff et al. (2023) found that machine-translating human text dropped correct “human” classifications from about 95% to 70%.

Insist on calibrated confidence, not binary verdicts. The survey’s own recommendation is pointed: knowing whether a system is 55% certain or 99% certain should change what you do next. Binary “AI / Human” labels invite exactly the overconfidence that the human-detection studies should have cured us of.

Design the workflow so high-confidence cases are automated and ambiguous cases go to a human – with the statistical evidence in front of them. That is the 54%-to-72% finding, operationalised.

Aggregate. More text, multiple documents, multiple signals, multiple tools. The survey’s conclusion is that ensemble approaches – several differently calibrated methods combined – are the most robust path currently available, and that submitting a text to multiple tools may give a more reliable answer than any one of them alone.

“In conclusion…” (we couldn’t resist)

Unaided human detection of AI-generated text performs between 50% and 61% across the studies in the literature. Familiarity with AI doesn’t reliably help. Linguistic training barely helps. Small amounts of practice help a little. What helps substantially – an 18-point jump – is putting statistical evidence in front of the human reader.

The stylistic tells everyone trades in are real patterns, but they are weak, genre-dependent, erased by a single editing pass, and systematically biased against non-native English writers. They are the worst possible foundation for a high-stakes accusation.

You probably can’t tell. Neither can we, by eye. That’s precisely why the measurement has to be done in a way that doesn’t depend on eyes – and why the honest version of that measurement reports its uncertainty instead of pretending to certainty it doesn’t have.

Try our AI detector now – register here.

Sources

  • Ardito, C. G. (2024). Contra generative AI detection in higher education assessments. New Directions for Teaching and Learning, special issue: Integrating Generative AI in the Design of Assessment.
  • Bao, G., Zhao, Y., Teng, Z., Yang, L., & Zhang, Y. (2024). Fast-DetectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR).
  • Chakraborty, M., Tonmoy, S. T. I., Zaman, S. M., Gautam, S., Kumar, T., Sharma, K., Barman, N., Gupta, C., Jain, V., Chadha, A., et al. (2023). Counter Turing Test (CT²): AI-generated text detection is not as easy as you may think — introducing AI Detectability Index (ADI). In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2206–2239).
  • Cui, W., Zhang, L., Wang, Q., & Cai, S. (2023). Who said that? Benchmarking social media AI detection. arXiv preprint arXiv:2310.08240.
  • Fraser, K. C., Dawkins, H., & Kiritchenko, S. (2025). Detecting AI-generated text: Factors influencing detectability with current methods. Journal of Artificial Intelligence Research, 82, 2233–2278. https://doi.org/10.48550/arXiv.2406.15583 (CC BY 4.0)
  • Gehrmann, S., Strobelt, H., & Rush, A. M. (2019). GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
  • Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., & Wu, Y. (2023). How close is ChatGPT to human experts? Comparison corpus, evaluation, and detection. arXiv preprint arXiv:2301.07597.
  • He, X., Shen, X., Chen, Z., Backes, M., & Zhang, Y. (2024). MGTBench: Benchmarking machine-generated text detection. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (pp. 2251–2265).
  • Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS).
  • Li, Y., Li, Q., Cui, L., Bi, W., Wang, Z., Wang, L., Yang, L., Shi, S., & Zhang, Y. (2024). MAGE: Machine-generated text detection in the wild. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 36–53). Bangkok, Thailand: ACL.
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779.
  • Liu, Y., Zhang, Z., Zhang, W., Yue, S., Zhao, X., Cheng, X., Zhang, Y., & Hu, H. (2023). ArguGPT: Evaluating, understanding and identifying argumentative essays generated by GPT models. arXiv preprint arXiv:2304.07666.
  • Liu, Z., Yao, Z., Li, F., & Luo, B. (2024). On the detectability of ChatGPT content: Benchmarking, methodology, and evaluation through the lens of academic writing. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (pp. 2236–2250).
  • Mazar, N., Amir, O., & Ariely, D. (2008). The dishonesty of honest people: A theory of self-concept maintenance. Journal of Marketing Research, 45(6), 633–644.
  • Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning (ICML).
  • Muñoz-Ortiz, A., Gómez-Rodríguez, C., & Vilares, D. (2023). Contrasting linguistic patterns in human and LLM-generated text. arXiv preprint arXiv:2308.09067.
  • Nguyen-Son, H.-Q., Tieu, N.-D. T., Nguyen, H. H., Yamagishi, J., & Zen, I. E. (2017). Identifying computer-generated text using statistical analysis. In Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (pp. 1504–1511).
  • Pagnoni, A., Graciarena, M., & Tsvetkov, Y. (2022). Threat scenarios and best practices to detect neural fake news. In Proceedings of the 29th International Conference on Computational Linguistics (COLING) (pp. 1233–1249).
  • Pu, J., Sarwar, Z., Abdullah, S. M., Rehman, A., Kim, Y., Bhattacharya, P., Javed, M., & Viswanath, B. (2023). Deepfake text detection: Limitations and opportunities. In Proceedings of the IEEE Symposium on Security and Privacy (SP) (pp. 1613–1630). IEEE.
  • Sarvazyan, A. M., González, J. Á., Franco-Salvador, M., Rangel, F., Chulvi, B., & Rosso, P. (2023). Overview of AuTexTification at IberLEF 2023: Detection and attribution of machine-generated text in multiple domains. Procesamiento del Lenguaje Natural, 71, 275–288.
  • Schuster, T., Schuster, R., Shah, D. J., & Barzilay, R. (2020). The limitations of stylometry for detecting machine-generated fake news. Computational Linguistics, 46(2), 499–510.
  • Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., et al. (2019). Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.
  • Spitale, G., Biller-Andorno, N., & Germani, F. (2023). AI model GPT-3 (dis)informs us better than humans. Science Advances, 9(26).
  • Stiff, H., & Johansson, F. (2022). Detecting computer-generated disinformation. International Journal of Data Science and Analytics, 13(4), 363–383.
  • Tian, E. (2023). Identifying GPT: First principles for generative AI detection (M.Sc. thesis). Princeton University.
  • Tulchinskii, E., Kuznetsov, K., Kushnareva, L., Cherniavskii, D., Nikolenko, S., Burnaev, E., Barannikov, S., & Piontkovskaya, I. (2023). Intrinsic dimension estimation for robust detection of AI-generated texts. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS) (pp. 39257–39276).
  • Uchendu, A., Ma, Z., Le, T., Zhang, R., & Lee, D. (2021). TuringBench: A benchmark environment for Turing test in the age of neural text generation. In Findings of the Association for Computational Linguistics: EMNLP 2021 (pp. 2001–2016).
  • Verma, V., Fleisig, E., Tomlin, N., & Klein, D. (2024). Ghostbuster: Detecting text ghostwritten by large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 1702–1717). Mexico City, Mexico: ACL.
  • Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1), 26.
  • Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending against neural fake news. Advances in Neural Information Processing Systems, 32.
  • Zhang, Q., Gao, C., Chen, D., Huang, Y., Huang, Y., Sun, Z., Zhang, S., Li, W., Fu, Z., Wan, Y., & Sun, L. (2024). LLM-as-a-coauthor: Can mixed human-written and machine-generated text be detected? In Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 409–436). Mexico City, Mexico: ACL.

Leave a Comment

Find us on: