SyloSpace

Do AI text detectors actually work?

Studies from 2023 to 2025 find AI-text detectors unreliable at telling human and AI writing apart, and advise against treating their output as proof.

Updated 15 hours ago6 min readVersion 4
CommentsFollow

Covers: Automated detectors for text generated by large language models, especially in education. Watermarking is mentioned only for context.

3 free full reads left this month. Join or upgrade

The short answer

Evidence-backed AI-organised, reviewed

The evidence does not support relying on AI-text detectors to tell human and AI writing apart. Two 2023 peer-reviewed evaluations found widely used detectors neither accurate nor reliable: one study of seven detectors misclassified over half of TOEFL essays by non-native English speakers as AI-generated (average false-positive rate 61%) while handling US eighth-grade essays almost perfectly, and a broader test of 12 public tools plus two commercial systems (Turnitin, PlagiarismCheck) found a main bias toward labelling text as human-written. Later work shows the picture is uneven rather than uniformly bad: a 2024 test of 10 free detectors found sensitivity ranging from 0% to 100%, with 5 of 10 detecting AI content at 100%, and a 2023 comparison of 16 detectors found three (Copyleaks, TurnItIn, Originality.ai) accurate across all document sets while most others failed to distinguish GPT-4 papers from undergraduate essays. A 2025 review of 34 articles concluded that despite most detectors exceeding 50% accuracy, they remain unreliable, with paid tools generally outperforming free ones.12345

What this rests on6 independent sources · 2 versions
  • Evidence 22
  • Interpretation 3

In brief

  1. Across multiple evaluations from 2023 to 2025, AI-text detectors are consistently described as unreliable for distinguishing human from AI writing.125

    Evidence-backed
  2. Seven detectors misclassified over half of TOEFL essays by non-native English speakers as AI-generated (average 61% false positives), while US eighth-grade essays were classified almost perfectly.1

    Evidence-backed
  3. Performance varies sharply by tool and text type: 5 of 10 free detectors reached 100% sensitivity in one 2024 test, while most of 16 detectors in a 2023 comparison failed to distinguish GPT-4 essays from undergraduate writing.34

    Evidence-backed
  4. Simple prompting such as asking for more literary language, and paraphrasing or other obfuscation, sharply reduce detection of AI-generated text.125

    Evidence-backed
  5. Reviews and the original studies advise against treating detector output as proof in educational or evaluative settings, and recommend pairing tools with human judgment.125

    Evidence-backed

At a glance

The picture in numbers

Live · updated just now

Average false-positive rate in a study of seven detectors

61%

61 in every 100

of TOEFL essays by non-native English speakers flagged as AI-written1
Test of 10 free AI-detector tools in 2024
  • detected AI at 100%5 of 10 tools
  • did not5 of 10 tools
Detectors that caught AI text at 100% sensitivity3
2025 literature review

34 articles

34 articles: articles reviewed on detector reliability5

The evidence behind it

6 sources
  • Reviews of many studies1
  • Other studies and data5

When it was published

Newest from 2025

20232026
Sources on this page by kind and year
SourceKindYear
GPT detectors are biased against non-native English writersOther studies and data2023
Testing of detection tools for AI-generated textOther studies and data2023
Enhancing the Robustness of AI-Generated Text Detectors: A SurveyOther studies and data2025
How Sensitive Are the Free AI-detector Tools in Detecting AI-generated Texts? A Comparison of Popular AI-detector ToolsOther studies and data2024
The Effectiveness of Software Designed to Detect AI-Generated Writing: A Comparison of 16 AI Text DetectorsOther studies and data2023
Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature ReviewReviews of many studies2025

The community around it

Contributions
0
People
0
Following
0

Nobody has added anything yet. Experience, evidence or a different view would show up here.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you are a teacher or institution considering detector output as evidence of misconduct

the documented error rates and the authors' caution against educational use argue for treating detector output as a weak signal, not proof.125

Interpretation

If your students include non-native English writers

expect elevated false-positive risk: over half of TOEFL essays in one study were flagged as AI-generated, and a 2025 review flags bias against non-native English speakers as an ongoing concern.15

Evidence-backed

If you are relying on a detector to catch AI-generated submissions

note that the broad tool test found a bias toward the human-written label, so AI text may pass undetected, and that paraphrasing or obfuscation worsens performance further.25

Evidence-backed

If you assume a student could not have evaded detection

simple prompting changes such as more literary language sharply reduced detection in one study, so evasion may not require technical skill.1

Evidence-backed

If you are choosing a detector for academic use

performance is tool-specific: three named detectors were accurate across GPT-3.5, GPT-4 and student essays in one comparison, and 5 of 10 free tools reached 100% sensitivity in another, while most other tested tools were unreliable, so no tested option is a dependable sole check.432

Evidence-backed

If you are weighing free versus paid detectors

a 2025 review found paid tools generally perform better than free ones, though a 2023 comparison found paid detectors only slightly more accurate than the rest.54

Evidence-backed

If you are evaluating a detector against newer model output

expect weaker results: most of 16 detectors that handled GPT-3.5 essays reasonably well were ineffective at distinguishing GPT-4 papers from undergraduate writing.4

Evidence-backed

If you need a defensible process rather than a verdict

the 2025 review advises using detectors only partially and combining them with human judgment to identify the true writer of a text.5

Evidence-backed

The full story · 3 chapters

01

What the tests found

AI summary:Multiple tests found detectors inaccurate and unreliable, with results varying widely by tool and text type.

Evidence-backed

Evidence-backed: In a study of seven widely used GPT detectors, over half of TOEFL essays written by non-native English speakers were classified as AI-generated, an average false-positive rate of 61%, while essays by US eighth-graders were classified almost perfectly. The same study found that simple prompting strategies, such as asking a model to use more literary language, sharply reduced detection of AI-generated text. The authors caution against using detectors in evaluative or educational settings.1

Evidence-backed

Evidence-backed: A separate test covering 12 publicly available tools and two commercial systems widely used in academia (Turnitin and PlagiarismCheck) concluded that available detection tools are neither accurate nor reliable, with a main bias toward classifying output as human-written rather than detecting AI-generated text. Content obfuscation techniques significantly worsened tool performance.2

Evidence-backed

Evidence-backed: A 2024 test of 10 free AI-detector tools found sensitivity ranging from 0% to 100%, with 5 of the 10 detecting AI-generated content at 100% accuracy. On paraphrased texts, Sapling and Undetectable AI identified all three software-generated contents at 100% accuracy, while Copyleaks, QuillBot and Wordtune identified content from two software programs at 100% accuracy.3

Evidence-backed

Evidence-backed: A comparison of 16 publicly available detectors tested on 42 ChatGPT-3.5 essays, 42 ChatGPT-4 essays and 42 student essays found three detectors — Copyleaks, TurnItIn and Originality.ai — highly accurate across all three sets. Most of the other 13 could distinguish GPT-3.5 papers from human writing with reasonably high accuracy but were generally ineffective at distinguishing GPT-4 papers from undergraduate essays. Detectors requiring registration and payment were only slightly more accurate than the others.4

Evidence-backed

Evidence-backed: A 2025 literature review of 34 articles found that although most detectors attained accuracy above 50%, they are unreliable; paid tools generally perform better than free ones, but there are concerns about bias against non-native English speakers, and the tools struggle with sophisticated AI content and tricks like paraphrasing.5

Participant opinion · poll

Where do you meet AI detectors?

Where do you meet AI detectors?As a studentAs a teacherAt workNot at all
Sign in to respond.

Your individual response is private. Only totals are shown.

02

Why detection is fragile

AI summary:Detectors flag human writing, miss AI writing, and are easily defeated by paraphrasing or simple prompting.

Interpretation

Interpretation: Several failure modes are documented. Detectors can flag human writing as AI: non-native English writing was misclassified at high rates, suggesting detectors may key on stylistic features that correlate with language background rather than with machine generation. Detectors can also miss AI writing: the broad tool test found a bias toward the human-written label, and content obfuscation made performance significantly worse. And detection is uneven across model generations: most of 16 detectors handled GPT-3.5 essays reasonably well but failed on GPT-4 essays.124

Evidence-backed

Evidence-backed: Detection also appears easy to defeat by ordinary means: asking a model to write in a more literary style sharply reduced detection of its output, and paraphrasing is repeatedly identified as a technique that degrades detector performance.135

Evidence-backed

Evidence-backed: A 2025 survey of detector robustness organises the weaknesses into three areas — text perturbation, out-of-distribution text, and AI–human hybrid text — and reports that all three affect the performance of commonly used detectors, leaving significant room for improvement.6

03

What this means for education

AI summary:The studies advise against using detector output as proof in academic settings, especially given bias against non-native writers.

Evidence-backed

Evidence-backed: The studies were conducted with academic settings in mind. One explicitly cautions against using detectors in evaluative or educational settings; another frames its findings around the implications and drawbacks of using detection tools in academic settings; and a 2025 review advises that authorities using these detectors should only partially trust them, since they are imperfect and can still misjudge, and that users should not rely completely on the tools but cooperate with them to find the true writer of a text.125

Interpretation

Interpretation: The practical risk is asymmetric by group: if a detector's errors fall disproportionately on non-native English writers, then using detector output as evidence of misconduct would penalise some students more than others. The 2025 review raises the same concern about bias against non-native English speakers.15

Evidence-backed

Evidence-backed: Tool choice matters more than the category 'detector' suggests: three named detectors were accurate across GPT-3.5, GPT-4 and student essays in one comparison, and 5 of 10 free tools reached 100% sensitivity in another test, while most other tools in the same tests performed poorly.43

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    GPT detectors are biased against non-native English writers
    Patterns (Liang et al.)Published Jul 1, 2023Checked Sep 30, 2026
    “Seven widely used GPT detectors misclassified over half of TOEFL essays written by non-native English speakers as AI-generated (an average false-positive rate of 61%), while classifying essays by US eighth-graders almost perfectly. Simple prompting strategies, such as asking a model to use more literary language, sharply reduced detection of AI-generated text. The authors caution against using detectors in evaluative or educational settings.”
  2. 2
    Testing of detection tools for AI-generated text
    International Journal for Educational Integrity (Weber-Wulff et al.)Published Dec 24, 2023Checked Sep 30, 2026
    “Specifically, the study seeks to answer research questions about whether existing detection tools can reliably differentiate between human-written text and ChatGPT-generated text, and whether machine translation and content obfuscation techniques affect the detection of AI-generated text. The research covers 12 publicly available tools and two commercial systems (Turnitin and PlagiarismCheck) that are widely used in the academic setting. The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools. The study makes several significant contributions. First, it summarises up-to-date similar scientific and non-scientific efforts in the field. Second, it presents the result of one of the most comprehensive tests conducted so far, based on a rigorous research methodology, an original document set, and a broad coverage of tools. Third, it discusses the implications and drawbacks of using detection tools for AI-generated text in academic settings.”
  3. 3
    How Sensitive Are the Free AI-detector Tools in Detecting AI-generated Texts? A Comparison of Popular AI-detector Tools
    Indian Journal of Psychological Medicine (Kar et al.)Published May 11, 2024Checked Oct 4, 2026
    “10 AI-detector tools were tested on their ability to detect AI-generated text. The sensitivity ranged from 0% to 100%. 5 out of 10 tools detected AI-generated content with 100% accuracy. For paraphrased texts, Sapling and Undetectable AI detected all three software-generated contents with 100% accuracy. Meanwhile, Copyleaks, QuillBot , and Wordtune identified content generated by two software programs with 100% accuracy. The integration of AI technology in academic writing is becoming more prevalent. Nonetheless, relying solely on AI-generated content can diminish the author’s credibility, leading most academic journals to suggest limiting its use. AI-content-detection software programs have been developed to detect AI-generated or AI-assisted texts. Currently, some of the platforms are equally sensitive.”
  4. 4
    The Effectiveness of Software Designed to Detect AI-Generated Writing: A Comparison of 16 AI Text Detectors
    Open Information Science (Walters)Published Jan 1, 2023Checked Oct 4, 2026
    “This study evaluates the accuracy of 16 publicly available AI text detectors in discriminating between AI-generated and human-generated writing. The evaluated documents include 42 undergraduate essays generated by ChatGPT-3.5, 42 generated by ChatGPT-4, and 42 written by students in a first-year composition course without the use of AI. Each detector’s performance was assessed with regard to its overall accuracy, its accuracy with each type of document, its decisiveness (the relative number of uncertain responses), the number of false positives (human-generated papers designated as AI by the detector), and the number of false negatives (AI-generated papers designated as human). Three detectors – Copyleaks, TurnItIn, and Originality.ai – have high accuracy with all three sets of documents. Although most of the other 13 detectors can distinguish between GPT-3.5 papers and human-generated papers with reasonably high accuracy, they are generally ineffective at distinguishing between GPT-4 papers and those written by undergraduate students. Overall, the detectors that require registration and payment are only slightly more accurate than the others.”
  5. 5
    Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review
    American Journal of IR 4 0 and Beyond (Gotoman et al.)Published Feb 18, 2025Checked Oct 4, 2026
    “This study used a literature review wherein pertinent studies were gathered and selected to discover potential implications. Research objectives were defined to assess the accuracy and reliability of the AI text detectors and identify which AI detectors were evaluated. Three online databases were used to search for relevant literature, of which 34 articles were finalized. Results show that despite most detectors attaining accuracy above 50%, they are unreliable. Paid tools generally perform better than free ones, but there are concerns about bias against non-native English speakers. These tools also struggle with sophisticated AI content and tricks like paraphrasing, so using them carefully and relying on human judgment is important to avoid unfairly discrediting someone’s work. AI-generated text detection technology still has a lot of room for improvement. Users should not rely completely on these tools but rather cooperate with those tools to better find the true writer of a text. Hence, authorities who use these AI detectors should only partially trust these tools, for they are imperfect and can still make mistakes in their judgment.”
  6. 6
    Enhancing the Robustness of AI-Generated Text Detectors: A Survey
    Mathematics (Liu et al.)Published Jun 30, 2025Checked Oct 4, 2026
    “This survey provides a systematic overview of existing research on enhancing the robustness of AIGT detectors. We categorize the focus of related literature into three key areas: text perturbation robustness, out-of-distribution (OOD) robustness, and AI–human hybrid text (AHT) detection robustness. For each area, we thoroughly summarize and analyze the corresponding robustness enhancement methods and additionally incorporate some approaches from other fields as a supplement. We also methodically organize relevant benchmark datasets, robustness evaluation methods, and metrics used to assess detectors’ performance. Then, through experiments, we evaluate the robustness of several commonly used detectors. Experiments show that text perturbations, OOD text, and AHT all affect the performance of these detectors, revealing that there remains significant room for improvement in their robustness. Finally, we suggest promising future directions based on the current issues faced by AIGT detectors and the detection requirements in real-world scenarios. To the best of our knowledge, this is the first review focused specifically on the robustness of AIGT detection.”

How it changed

Published 2 times since Sep 30, 2026.

  1. Version 4Oct 4, 2026Live now

    Added three newer sources (2024–2025) that broaden the picture beyond the two 2023 evaluations: a 2024 test of 10 free detectors showing sensitivity from 0% to 100%, a 2023 comparison of 16 detectors finding most fail on GPT-4 text, and a 2025 literature review of 34 articles concluding detectors remain unreliable despite often exceeding 50% accuracy. Added a robustness survey documenting perturbation, out-of-distribution and hybrid-text weaknesses.

    • The main finding was rewritten.
    • Updated “What the tests found”.
    • Updated “Why detection is fragile”.
  2. Version 3Sep 30, 2026

    Initial Starting Map on whether AI-text detectors can reliably distinguish human from AI writing, built from two peer-reviewed evaluations covering 2023.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

Open questions

  • How do detectors released or updated after these studies perform on the same kinds of texts, especially against newer model outputs?

    No answers yet

  • Across tools and text types, which error is more common in practice: flagging human writing as AI, or missing AI writing?

    No answers yet

  • What mechanisms cause higher false-positive rates for non-native English writers, and can they be corrected?

    No answers yet

  • What happens to students accused on the basis of detector output, and what safeguards do institutions use?

    No answers yet

  • How well do detectors handle text that mixes human and AI writing, and does that change the accuracy picture?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Add what you know

Sign in to add what you know. Reading stays open to everyone.

Ask this Sylo

Answers only from “Do AI text detectors actually work?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.