arXiv:2605.09228cs.LGcs.AI2026-05

评测大模型主动发现用户隐含需求的能力,填补传统评测空白。

ProactBench: Beyond What The User Asked For

论文配图:ProactBench: Beyond What The User Asked For
图 1 · 摘自论文原文
  • 将对话主动性拆解为三种可测阶段:突发、综合与任务后恢复。
  • 在16个模型中,'任务后恢复'能力最弱且现有基准无法有效预测。
  • 构建了含624个触发点的独立验证数据集,适合评估高级对话能力。

多数大模型评测仅关注模型对明确指令的响应能力,却忽视了一种更深层的对话能力:察觉并回应用户未明说但隐含的需求,我们称之为‘对话主动性’。本文提出ProactBench,将该能力分解为三类时序相关的类型: extsc{Emergent}(基于单一显性线索推断)、 extsc{Critical}(跨多个线索综合判断)和 extsc{Recovery}(任务完成后的前瞻价值维持)。通过设计规划者、用户代理与助手模型三者间的信息不对称,有效规避风格混淆、评分标准泄露、外部上下文污染与信息过载等问题。释放的数据集包含198段精心筛选的对话,涵盖624个触发点,覆盖24种由心理量表生成的沟通风格,并经独立大模型裁判审计。在16个前沿及开源模型上测试显示, extsc{Recovery}既困难又难以被六个主流基准预测,因此成为极具价值的新评估信号。

原文摘要 · Abstract (English)

Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implied but not said. We call this \emph{conversational proactivity}. ProactBench decomposes it into three phase-tied types: \textsc{Emergent}, inference from a single disclosed anchor; \textsc{Critical}, synthesis across multiple anchors; and \textsc{Recovery}, grounded forward-looking value after task completion. We operationalise the benchmark with three agents: a Planner, a User Agent, and an Assistant Model. Their information asymmetries defend against style-confounded scoring, rubric leakage, external-context contamination, and information dumps. The released corpus contains 198 curated dialogues with 624 trigger points across 24 communication styles drawn from a psychometric inventory and audited by an independent LLM judge. Across 16 frontier and open-weight models, \textsc{Recovery} is both difficult and weakly predicted by six standard benchmarks, making it a useful new evaluation signal.

对话评估大模型评测主动对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。