研究推理表示对人类评估大模型输出的帮助,发现好看不等于好用。
Do Reasoning Representations Help Humans Evaluate LLM Outputs?

- 把推理过程当人看的界面,测试六种格式在不同任务中的表现。
- 简单思维链更利于验证、信任和可解释性,但人更喜欢复杂格式。
- 偏好复杂格式易导致误判,高信任却不愿核对,有风险。
推理表示被越来越多地用作大语言模型输出的解释,但其评估通常基于模型中心指标(如答案准确率和忠实度),未明确其是否真正帮助人类评估模型响应。本文将推理表示视为面向人类的交互界面,而非模型推理能力的代理。我们通过一个受控的人类实验,在不同复杂度的任务中对比六种推理格式,依托基于网页的框架随机化任务领域、问题实例和表示顺序。实验收集了关于结构理解、错误检测与定位、信任校准的细粒度判断。结果表明,人类偏好与实际支持存在错位:参与者更青睐基于规划和分解的表示形式,但更简单的思维链(Chain-of-Thought)痕迹在支持验证、信任建立和可解释性方面表现更优。此外,偏好格式带来校准风险——对正确轨迹产生更多误报,且即使不愿验证也表现出高信任度。
原文摘要 · Abstract (English)
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。