{"ok":true,"article":{"slug":"why-ai-detection-fails-for-academic-integrity-185e85ae","title":"Why AI Detection Fails for Academic Integrity","url":"https://arxiv.org/abs/2608.11256","canonical":"https://www.aimode.news/article/why-ai-detection-fails-for-academic-integrity-185e85ae","sourceName":"arXiv cs.LG","summary":"arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light \"refine abstract only\" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.","category":"AI","image":null,"lang":"en","publishedAt":"2026-08-13T04:00:00+00:00","createdAt":"2026-08-13T04:30:12.663308+00:00"}}