测试大模型在信息不足时主动提问的能力,发现它们缺乏真正智能的主动性。
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- 构建两类不完整问题数据集,模拟真实场景中的信息缺失。
- 多数大模型无法主动请求补充信息,仅能强行作答。
- 适合关注模型推理可靠性与人机交互的研究者。
大型推理模型(LRMs)在数学问题求解方面表现卓越,但现有评估基准仅针对定义明确的问题,存在关键缺陷:真正智能的代理不仅应能解题(如数学测验答题),还应在信息不足时主动请求补充信息,展现对用户需求的主动性响应。为填补这一空白,本文提出一个包含两类不完整问题的新数据集,涵盖多样情境。基于该数据集的系统性评估揭示,当前LRMs普遍缺乏主动询问信息的能力。此外,研究还发现了模型存在的过度思考与幻觉行为,并探讨了监督微调在学习此类能力上的潜力与挑战。本工作旨在推动开发具备真正智能的模型,而非仅限于解题。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have demonstrated remarkable problem-solving abilities in mathematics, as evaluated by existing benchmarks exclusively on well-defined problems. However, such evaluation setup constitutes a critical gap, since a genuine intelligent agent should not only solve problems (as a math quiz solver), but also be able~to ask for information when the problems lack sufficient information, enabling proactivity in responding users' requests. To bridge such gap, we proposes a new dataset consisting of two types of incomplete problems with diverse contexts. Based on the dataset, our systematical evaluation of LRMs reveals their inability in proactively asking for information. In addition, we uncover the behaviors related to overthinking and hallucination of LRMs, and highlight the potential and challenges of supervised fine-tuning in learning such ability. We hope to provide new insights in developing LRMs with genuine intelligence, rather than just solving problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。