用工具评估大模型下棋解说,发现幻觉普遍且难根治。
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

- 拆解棋局解说为原子命题,结合引擎与专家标注验证真伪
- 无工具时大模型错误率高达40%,工具提升后仍难覆盖战略思想
- 适合研究模型可信度、教育类应用或可解释性评估的开发者
超人级棋类引擎虽能提供专家级判断,但缺乏自然语言解释,难以用于教学。大语言模型理论上可填补此空白,却因领域知识有限而常产生幻觉,现有评估框架无法可靠检测。本文提出ACT-Eval,将棋局解说分解为原子命题,通过引擎工具与专家标注的黄金参考进行事实正确性、概念覆盖度和走法质量评估。我们发布包含325个对局-着法对的基准数据集,涵盖教学、赛事与关键局面,含125个经专家验证的黄金原子及五类错误分类。评估主流闭源与开源模型发现:未使用工具时,GPT-5.4的子命题错误率达22.0%,小模型超过40%;工具增强虽显著提升事实准确性与走法评估,但对专家级战略战术思想的覆盖仍严重不足。人类校准显示,ACT-Eval的事实判断在人际一致性范围内,覆盖度评分与人类对战略完整性的评价高度相关。
原文摘要 · Abstract (English)
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but they frequently hallucinate due to limited domain-specific knowledge, and standard reference-based or LLM-as-a-judge frameworks cannot reliably detect these errors. In this work, we present ACT-Eval, an evaluation framework that decomposes chess commentary into atomic claims and routes them to engine-supported tools and expert-annotated gold references to assess factual correctness, conceptual coverage, and move-quality judgment. We release a benchmark of 325 position--move pairs spanning pedagogical, tournament, and critical positions, including 125 positions with expert-verified gold atoms and a five-class error taxonomy. Evaluating leading proprietary and open-weight models, we find that factual hallucinations remain pervasive in chess commentary: GPT-5.4 without tools produces incorrect sub-claims 22.0% of the time, while smaller open-weight models exceed 40%. Although tool augmentation substantially improves factual correctness and move-quality assessment, coverage of expert strategic and tactical ideas remains limited across all models. Human calibration shows that ACT-Eval's factual judgments fall within the observed range of inter-human agreement, while its coverage scores correlate strongly with human assessments of strategic completeness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。