为地图智能体设计新评测基准,聚焦用户未明说的隐性需求。
MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors

- 从用户行为链中还原完整需求,识别可推断的隐性决策因素。
- 在真实数据上构建多维度标注基准,量化评估满意度相关能力。
- 揭示现有智能体对隐性需求响应不足,适合空间决策研究者使用。
大型语言模型代理正被越来越多地集成到地图服务中。由于地图服务嵌入日常场景而非专业任务环境,用户常以非正式方式表达需求,导致查询信息不全,存在大量未言明的隐性决策因素,这些因素对用户满意度至关重要。尽管澄清是缓解此问题的有效手段,但会增加用户负担;因此,一个高效的代理应能主动从已有信息源中恢复这些因素。然而,评估该能力面临双重挑战:一是需确定哪些隐性因素可被评估——仅当其影响用户接受度且代理可在回应前获取相关证据时才具备可评估性;二是用户满意度无法由单一参考答案代表,必须将满意度相关因素转化为客观可量化的评价目标。为此,我们提出一个‘还原-识别-过滤’框架,通过行为链证据重建完整用户需求,识别隐性决策因素,并保留仅由预查询证据支持的因素。基于该方法,我们利用大规模真实匿名用户数据构建了MapSatisfyBench,从五个维度标注真实情况,支持对满意度感知地图代理的全流程评估。实验表明,当前代理在显式任务完成上表现良好,但在满足隐性决策因素及主动获取满意度所需证据方面仍存局限。该成果确立了MapSatisfyBench作为推动地图代理评估从任务完成转向满意度感知空间决策的新基准。
原文摘要 · Abstract (English)
Large language model agents are increasingly integrated into map services. Since map services are embedded in everyday-life scenarios rather than professional task settings, users often express their needs informally, resulting in underspecified queries with many unspoken needs, namely, implicit decision factors that are critical for user satisfaction. Although clarification is an effective way to mitigate this issue, it increases user burden in daily interaction, and a capable agent should first proactively recover such factors from available information sources. However, evaluating this ability is challenging. The first challenge is to determine which implicit decision factors are suitable for evaluation. A factor is evaluable only if it affects user acceptance and can be recovered from information available to the agent before it responds. Second, user satisfaction cannot be reliably represented by a single reference answer, requiring a benchmark that converts satisfaction-relevant factors into objective and quantifiable evaluation targets. To address these challenges, we propose a restore-identify-filter framework that reconstructs complete user needs from behavior-chain evidence, identifies implicit decision factors, and retains only those supported by pre-query evidence. Building on this methodology, we construct MapSatisfyBench from large-scale, real-world anonymized user data and annotate ground truth from five dimensions and enables full-chain evaluation of satisfaction-aware map agents. Experiments show that current agents generally perform well on explicit task completion, but remain limited in satisfying implicit decision factors and proactively acquiring the evidence needed for satisfaction-aware decisions. These findings establish MapSatisfyBench as a benchmark for shifting map-agent evaluation from task completion toward satisfaction-aware spatial decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。