Epoch AI Finds Detectors Miss 26% of Style-Mimicked Scientific Writing
Epoch AI stress-tested three major AI detectors (Pangram, GPTZero, Originality.ai) against 99 authors across real human writing, basic AI prompts, and AI text mimicking a specific author's style. The detectors caught nearly everything from simple prompts but missed roughly 13% of style-mimicked passages overall โ and 26% when the target was scientific writing.
Why it matters
๐ป Developer ยท If you're building or relying on AI-detection tooling, this is a concrete failure mode to test for โ style-mimicry attacks meaningfully degrade detector accuracy versus naive prompts.
๐ฆ Product ยท Any product that relies on AI-content detection as a trust or moderation feature should treat this as a known blind spot, not a solved problem.
๐จ Design ยท No direct design impact โ this is a research and policy story.
๐ Business ยท Academic and publishing institutions relying on these detectors for integrity enforcement are significantly overestimating their effectiveness โ worth flagging if your organization depends on AI detection for any compliance purpose.
๐ค Just Curious ยท Tools that claim to detect AI-written text turn out to miss a lot more than people think โ especially when the AI is specifically told to write in a particular person's style, which fools the detectors over a quarter of the time for scientific writing.