arXiv:2609.07784cs.AI2026-09

测试大模型解决真实生活咨询问题的能力,发现它们对隐含需求理解不足。

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

论文配图:xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
图 1 · 摘自论文原文
  • 构建248个真实场景任务,覆盖51种日常情境。
  • 顶尖模型任务完成率75.6%,隐含需求得分低9个百分点以上。
  • 强调模型需理解用户背景与未明说需求,适合做智能助手评估。

大型语言模型(LLMs)越来越多地用于日常生活辅助,但现有评测基准未能充分反映用户实际提出的请求。真实请求常为开放式、非正式描述且依赖上下文,要求模型不仅遵循显式指令,还需从用户背景和情境中推断未明说的需求。我们提出xDailyBench,一个包含248个精心设计任务的基准,覆盖个人生活、白领工作、学习科研及跨领域活动共51个场景。任务基于用户真实完成或意图使用AI实现的请求,采用细粒度二元评分标准,涵盖显性和隐性需求。我们在标准化代理设置下评估11个前沿模型,最优模型任务级得分为75.6%,所有模型在隐性需求上的表现均显著低于显性需求,差距不低于9个百分点。结果表明,隐性需求推理仍是满足真实世界用户日常需求的持续瓶颈。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.

大模型评测隐性需求日常咨询基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。