测试大模型能否不靠提示就识别职场中的深层问题结构。
KWBench: Measuring Unprompted Problem Recognition in Knowledge Work

- 设计223个真实职业场景任务,考察模型从原始信息中自主识别问题类型。
- 最佳模型仅正确识别27.9%的任务,顶尖模型间共识率不足32%。
- 强调问题识别能力,适合评估模型在复杂职场场景中的直觉判断力。
我们推出首个知识工作领域无提示问题识别基准KWBench:评估大语言模型能否在未被告知问题类型的情况下,从原始输入中识别出专业情境的本质结构。现有前沿评测已趋于饱和,多数知识工作评估聚焦于信息提取或按说明完成任务。KWBench则针对前置步骤——仅凭原始数据识别情境的支配结构。该基准包含223个来自并购、合同谈判、临床药学、组织政治、欺诈分析和激励设计等领域的任务,每项均对应一个形式化的博弈论模式(如委托代理冲突、信号传递、机制设计失败、策略性遗漏、联盟动态、战略相互依赖),并配有专家标注的情境解读与预期失效模式。模型仅接收原始数据与任务提示,无问题类型提示。评分采用三级嵌套标准,强制包含错误路径预测。评估16个模型,最优模型通过率为27.9%,前两名模型一致通过率仅31.7%;在前8名中,44个任务仅被单一模型解决,综合路由覆盖率达50.7%,接近单模型最优的两倍。一旦通过,质量评分趋同(约83%);但整体得分无收敛。相同模型在被明确提问时能准确说出博弈论概念,却无法在无提示下应用。KWBench旨在推动前沿模型评估范式转变,从‘问题已框定后执行’转向‘从情境中自主识别问题’。
原文摘要 · Abstract (English)
We introduce the first version of KWBench (Knowledge Work Bench), a benchmark for unprompted problem recognition in large language models: can an LLM identify a professional scenario before attempting to solve it. Existing frontier benchmarks have saturated, and most knowledge-work evaluations to date reduce to extraction or task completion against a specification. KWBench targets the step before that: recognizing the governing structure of the situation from raw inputs alone. The benchmark contains 223 tasks sourced from practitioners across acquisitions, contract negotiations, clinical pharmacy, organizational politics, fraud analysis, and incentive design. Each task encodes a formal game-theoretic pattern (principal-agent conflict, signaling, mechanism design failure, strategic omission, coalitional dynamics, strategic interdependence) and carries structured ground truth recording the expert reading of the situation and the anticipated failure modes. Models receive raw data and a task prompt with no indication of problem type. Scoring is a three-tier rubric gated by a mandatory conjunctive check. Mandatory criteria encode the predicted wrong paths. We evaluate 16 models. The best model passes on 27.9% of tasks. The top two models agree on only 31.7% of their passes. Among the top 8, 44 tasks are solved by exactly one model; routing across the top 8 covers 50.7% of the benchmark, nearly double the best single model. Conditional on passing, quality scores converge (approx 83% across models); unconditional scores do not. Same models articulate the relevant game-theoretic concept correctly when asked, then fail to apply it unprompted. We release KWBench to shift how frontier models are evaluated on knowledge work, scoring them on whether they recognize the right problem from the situation alone, not only on how well they execute once the problem has been framed for them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。