大模型自检生成文本能力有限,解释常出错。
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
- 让大模型判断自己或他人生成的文本,自检比互检更准
- 引入'不确定'类别后,检测准确率和解释质量均提升
- 适合关注大模型可信度与自我解释能力的研究者
随着大模型滥用风险上升,区分人类与大模型生成文本至关重要。本文在二分类(人类对大模型生成)和三分类(加入‘不确定’类)两种场景下,评估了6个大小不一的闭源与开源大模型的检测与解释能力。结果发现,大模型识别自身输出的表现优于识别其他模型输出,但整体仍不理想。引入三分类框架后,所有模型的检测准确率与解释质量均有提升。基于人工标注数据集的定量与定性分析揭示,主要解释失败源于错误特征依赖、幻觉及推理缺陷。研究凸显当前大模型在自检测与自解释方面的局限性,强调需进一步解决过拟合问题并提升泛化能力。
原文摘要 · Abstract (English)
Distinguishing between human- and LLM-generated texts is crucial given the risks associated with misuse of LLMs. This paper investigates detection and explanation capabilities of current LLMs across two settings: binary (human vs. LLM-generated) and ternary classification (including an ``undecided'' class). We evaluate 6 close- and open-source LLMs of varying sizes and find that self-detection (LLMs identifying their own outputs) consistently outperforms cross-detection (identifying outputs from other LLMs), though both remain suboptimal. Introducing a ternary classification framework improves both detection accuracy and explanation quality across all models. Through comprehensive quantitative and qualitative analyses using our human-annotated dataset, we identify key explanation failures, primarily reliance on inaccurate features, hallucinations, and flawed reasoning. Our findings underscore the limitations of current LLMs in self-detection and self-explanation, highlighting the need for further research to address overfitting and enhance generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。