用对抗游戏训练大模型发现隐藏错误,提升诊断能力。
Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis
- 构建对抗框架:一个角色造隐蔽错题,另一个角色找漏洞。
- 实测诊断准确率比GPT-4o高16.8%至31.4%。
- 适合研究模型鲁棒性与错误检测的学者使用。
大型语言模型在多个领域具备出色的推理与生成能力,但仍难以识别和诊断复杂错误。这主要源于训练目标侧重正确答案,限制了对错误的接触与学习。尽管近期研究引入错误信号,但多依赖浅层、静态错误,难以提升深层诊断能力。为此,我们提出隐匿与搜寻游戏(Hide and Seek Game, HSG),一种动态对抗式错误生成与诊断框架,并在数学问题求解任务上进行评估。HSG包含两个对抗角色:'隐匿者'(Sneaky)生成细微且具有欺骗性的推理错误,'诊断者'(Diagnosis)则试图精准检测这些错误。通过对抗共进化,双方的错误隐蔽性与诊断精度均得到提升。在多个数学推理任务上的实验表明,HSG显著增强了错误诊断能力,相较GPT-4o等基线模型,准确率提升16.8%至31.4%。我们还公开了一个包含欺骗性错误及其诊断标注的挑战性数据集,为未来研究提供基准。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in reasoning and generation across domains, but still struggle with identifying and diagnosing complex errors. This stems mainly from training objectives that prioritize correct answers, limiting exposure to and learning from errors. While recent studies have begun to address this by introducing error signals, most rely on shallow, static errors, restricting improvement in deep diagnostic ability. To overcome this, we propose Hide and Seek Game (HSG), a dynamic adversarial framework for error generation and diagnosis, and evaluate it on mathematical problem-solving. HSG involves two adversarial roles: Sneaky, which "hides" by generating subtle, deceptive reasoning errors, and Diagnosis, which "seeks" to accurately detect them. Through adversarial co-evolution, both error stealth and diagnostic precision are enhanced. Experiments on several math reasoning tasks show that HSG significantly boosts error diagnosis, achieving 16.8\%--31.4\% higher accuracy than baselines like GPT-4o. We also release a challenging dataset of deceptive errors and diagnostic annotations as a benchmark for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。