AI content detectors are facing a credibility test in 2025. After a year of rapid improvements, a new round of independent testing shows that accuracy has slipped for several standalone tools, even as general-purpose chatbots have emerged as surprisingly capable alternatives. The results matter for teachers, editors, publishers, and anyone trying to determine whether a piece of text was written by a human or generated by a machine.
Key Facts
- Only three standalone AI content detectors achieved perfect scores in the latest round: Pangram, QuillBot, and ZeroGPT.
- Several previously top-performing detectors declined, including Copyleaks, Originality.ai, and Undetectable.ai.
- Chatbots outperformed dedicated detectors: ChatGPT Plus, Copilot, and Gemini all scored perfectly on a five-text test.
- The ChatGPT free tier identified the writer of one human-written sample, raising privacy and identification concerns.
- Grok struggled, getting three of five classifications wrong and labeling all samples as human-written.
- False positives remain a serious problem, especially for non-native English writers.
- Plagiarism definitions apply to uncredited AI-generated text, even if no source was directly copied.
Why AI content detection is harder than it looks
The rise of generative AI since late 2022 has created a new kind of plagiarism problem. Merriam-Webster defines plagiarize as to steal and pass off the ideas or words of another as one's own, or to use another's production without crediting the source. That definition fits AI-generated content when a user presents machine-written text as their own without disclosure. The writer may not have copied a specific source, but the words were not theirs, and the true author, the AI model, goes uncredited.
Detecting that kind of text is not a simple technical task. AI detectors look for statistical patterns such as perplexity, burstiness, and token predictability. Human writing tends to be less uniform, with unexpected word choices and uneven sentence lengths. AI writing often is smoother and more predictable. But these signals overlap. A human editor can produce polished, formulaic prose. An AI model can be prompted to write with deliberate errors or varied rhythm. The result is an arms race between detectors and tools that claim to humanize AI text.
To test the current state of the art, a long-running independent evaluation used five blocks of text. Two were written by a human, and three were generated by ChatGPT. Each block was fed separately to each detector. A correct classification counted as a pass; an incorrect one counted as a failure. When a detector returned a percentage, anything above 70 percent was treated as a strong probability and taken as the detector's answer. The same five blocks were used across 11 standalone detectors, producing 55 individual tests.
The tested detectors included BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Monica, Originality.ai, QuillBot, Undetectable.ai, Writer.com, ZeroGPT, and a newcomer called Pangram. Monica was dropped because it limited free tests to 250 words and then required a $200 upgrade to continue. Pangram was added and immediately became one of the top performers. The test series has now been run six times since early 2023, allowing some comparison over time.
Overall standalone detector results
The latest results were mixed. Only three detectors earned a perfect score: Pangram, QuillBot, and ZeroGPT. Several others landed at 80 percent: Copyleaks, GPTZero, and Originality.ai. GPT-2 Output Detector scored 60 percent. BrandWell, Grammarly, and Writer.com each scored 40 percent. Undetectable.ai had the weakest showing at 20 percent.
The declines were notable. In earlier rounds, as many as five detectors had perfect scores. In April 2025, Copyleaks, QuillBot, Undetectable.ai, ZeroGPT, and others were performing at or near the top. By this latest round, Copyleaks and Originality.ai had fallen. Both flagged human-written text as AI-generated. Copyleaks had publicly claimed 99 percent accuracy backed by third-party studies, yet it identified a human-written sample as 100 percent AI. Originality.ai, which sells usage credits and describes itself as a most accurate AI detector, also labeled the human sample as AI-written.
Other tools showed different problems. GPTZero, which has grown from a bare-bones project into a company with a mission of protecting what is human, swapped its errors from one round to the next. In April, it got the first human sample wrong and the second AI sample right. In the latest test, it got the first human sample right and the second AI sample wrong. That kind of instability is difficult for users to trust.
Grammarly, a well-known writing assistant, showed no improvement in AI detection. Its tool labeled an entirely ChatGPT-written sample as human. Writer.com did the same for every block, even though three of the five samples were machine-generated. Undetectable.ai, which also offers a service to humanize AI text, rated a human-written sample as 60 percent likely AI and rated three AI-written samples as 75 to 77 percent likely human. Those results were the opposite of accurate.
On the positive side, ZeroGPT has matured. Early versions looked sketchy, with no clear company name and lots of ads. Now it presents as a typical software service with pricing, contact information, and a professional interface. Its accuracy rose from 80 percent to 100 percent and held there. QuillBot, which once gave wildly inconsistent results across multiple passes, has now achieved perfect scores in consecutive rounds. Pangram, founded by former engineers from major tech companies, offers five free tests per day. Its processing was slow, with a partially white screen between scan and result, but it scored five out of five.
The test series also looked for a broader trend over time. So far, there is no strong pattern of consistent improvement. Test 5 was reliably identified as human across detectors and dates, but even that reliability declined in the latest run. The only safe conclusion is that detector performance remains volatile.
Chatbots as content detectors
The latest round added a new experiment: using general-purpose chatbots as content detectors. Each chatbot was given the same prompt, followed by the text to check: Evaluate the following and tell me if it was written by a human or an AI
Source: ZDNET News