Skip to content

Can detectors reliably flag AI-written application answers?

Last updated 2026-08-17 · Confidence: documented — Pangram’s published numbers and a June 2026 independent evaluation; vendor accuracy figures are the vendor’s own.

Yes for text the AI mostly wrote: one detector, Pangram, now survives independent testing with near-zero false positives. The signal weakens as the human share of the writing grows, so a flag works as a conversation starter — never an auto-reject.

Pangram claims 99.98% accuracy and a 0.01% false-positive rate — about 1 in 10,000 human documents flagged. Its technical report reports error rates ~38x lower than commercial rivals and zero false positives on a non-native-English (TOEFL) benchmark — the classic detector failure.

Independent checks broadly agree. A June 2026 study in the International Journal for Educational Integrity tested four tools on 160 long academic papers: Pangram caught 97.5% of fully AI-generated papers and 95% of “humanised” ones, with zero false positives on human ESL writing — while GPTZero, Turnitin, and Copyleaks caught 0% of the fully-AI set. Pangram’s own third-party-evals roundup adds a University of Chicago audit (0.1% false positives) — read that page knowing the vendor curated it.

Mixed authorship is the hard case. Human edits obscure the statistical fingerprint, authorship flips between neighboring sentences, and short segments carry too few cues — the documented reasons hybrid text defeats detectors (GPTZero scored 0% on the hybrid papers above). Text a human drafted, dictated, and merely polished with AI is the weakest-signal case of all. Pangram labels segments human-written, AI-assisted, or AI-generated rather than giving one score, but the mostly-human end of that scale is where its evidence is thinnest.

Two cautions for screening applicants:

  • Pangram itself advises against scanning very short responses, bullet lists, and formulaic text — much of a typical form.
  • The IJEI authors’ conclusion: detectors “should not be used as sole evidence in high-stakes decision-making”.

So: flag, then weigh alongside the rest of the application — the same separate-evaluator discipline as testing LLM graders.