arXiv:2602.20424cs.AI2026-02

测试AI能否理解用户没说出口的需求,发现当前模型表现远未达人类水平。

Implicit Intelligence -- Evaluating Agents on What Users Don't Say

  • 用可读YAML定义交互世界,让模型通过探索发现隐藏约束。
  • 16个主流模型在205个场景中平均仅48.3%通过率,差距显著。
  • 适合关注具身智能、隐式推理与人机协作的研究者。

现实中的AI请求本质上是不明确的。自然语言交流依赖共享背景和未明说的约束,说话者期望听者能推断。现有代理评估基准仅测试显式指令遵循能力,无法评估模型是否能推理隐含需求,如无障碍需求、隐私边界、灾难性风险及上下文限制。我们提出隐式智能(Implicit Intelligence)评估框架,结合代理即世界(Agent-as-a-World, AaW),通过人类可读的YAML文件定义交互环境,并由语言模型模拟。场景设计为用户请求表面简单,实际解法复杂,需通过环境探索发现约束。在205个场景中评估16个前沿及开源模型,结果显示最优秀模型仅达48.3%场景通过率,揭示了从字面指令遵循到类人上下文推理之间仍有巨大提升空间。

原文摘要 · Abstract (English)

Real-world requests to AI agents are fundamentally underspecified. Natural human communication relies on shared context and unstated constraints that speakers expect listeners to infer. Current agentic benchmarks test explicit instruction-following but fail to evaluate whether agents can reason about implicit requirements spanning accessibility needs, privacy boundaries, catastrophic risks, and contextual constraints. We present Implicit Intelligence, an evaluation framework testing whether AI agents can move beyond prompt-following to become genuine goal-fulfillers, paired with Agent-as-a-World (AaW), a harness where interactive worlds are defined in human-readable YAML files and simulated by language models. Our scenarios feature apparent simplicity in user requests, hidden complexity in correct solutions, and discoverability of constraints through environmental exploration. Evaluating 16 frontier and open-weight models across 205 scenarios, we find that even the best-performing model achieves only 48.3% scenario pass rate, revealing substantial room for improvement in bridging the gap between literal instruction-following and human-like contextual reasoning.

隐式推理智能体评估上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。