用交互式问答评估模型推理能力,更真实反映智能水平。
Interactive Benchmarks
- 通过有限轮次交互评测模型获取与使用信息的能力
- 在逻辑、数学等任务中暴露现有模型的交互短板
- 适合关注模型决策与策略能力的研究者参考
现有推理评估方法存在局限:固定基准集趋于饱和且易受污染,基于偏好评估则依赖主观判断。我们认为,智能的核心在于决定获取何种信息并有效利用。为此提出交互式基准(Interactive Benchmarks),通过预算约束的多轮交互来评估模型推理能力。我们在两种场景下进行测试:交互证明(Interactive Proofs),模型与裁判互动解决逻辑、UI2Html和数学任务,获得客观反馈;交互游戏(Interactive Games),模型需制定策略以最大化长期收益。结果表明,该框架能更稳健地评估这一智能维度,揭示当前模型在交互场景中仍有巨大提升空间。
原文摘要 · Abstract (English)
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evaluations rely on subjective judgments. We argue that a core aspect of intelligence is the ability to decide what information to acquire and how to use it effectively. We propose Interactive Benchmarks, a unified evaluation paradigm that assesses a model's reasoning ability through budgeted multi-turn interaction. We evaluate models under this framework in two settings: Interactive Proofs, where models interact with a judge to solve Logic, UI2Html, and Mathematics tasks under objective feedback; and Interactive Games, where models reason strategically to maximize long-horizon utilities. Our results show that interactive benchmarks provide a more robust assessment of this dimension of model intelligence, revealing substantial room for improvement in interactive scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。