arXiv:2510.13499cs.CLcs.AI2025-10被引 2

首个面向真实消费讨论的动态意图理解评测基准,评估大模型在复杂对话中的综合推理能力。

ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding

  • 构建动态实时更新的消费者讨论数据集,模拟真实多视角交互场景。
  • 支持大规模、多样化的真实用户对话评测,覆盖多领域消费议题。
  • 适合评估大模型在不确定环境下整合信息与判断意图的能力,适用于产品设计与AI评测。

理解人类意图是大型语言模型面临的复杂高阶任务,需要分析推理、上下文解读、动态信息聚合及不确定性下的决策能力。现实世界中的公共讨论(如消费产品讨论)通常非线性且涉及多方观点,常存在交织冲突的立场、不同诉求、情感倾向以及隐含使用背景知识。要准确理解此类显性意图,大模型必须超越单句解析,整合多源信号,推理矛盾,并适应话语演变,如同政治、经济或金融领域的专家应对复杂环境。尽管该能力至关重要,但目前尚无大规模基准用于评估大模型在真实意图理解上的表现,主要受限于真实公共讨论数据的采集难度与评测流程构建挑战。为此,我们提出 ench,首个专为意图理解设计的动态、实时评测基准,尤其聚焦消费领域。ench 是同类中规模最大、多样性最高的基准,支持实时更新,并通过自动化清洗流程防止数据污染。

原文摘要 · Abstract (English)

Understanding human intent is a complex, high-level task for large language models (LLMs), requiring analytical reasoning, contextual interpretation, dynamic information aggregation, and decision-making under uncertainty. Real-world public discussions, such as consumer product discussions, are rarely linear or involve a single user. Instead, they are characterized by interwoven and often conflicting perspectives, divergent concerns, goals, emotional tendencies, as well as implicit assumptions and background knowledge about usage scenarios. To accurately understand such explicit public intent, an LLM must go beyond parsing individual sentences; it must integrate multi-source signals, reason over inconsistencies, and adapt to evolving discourse, similar to how experts in fields like politics, economics, or finance approach complex, uncertain environments. Despite the importance of this capability, no large-scale benchmark currently exists for evaluating LLMs on real-world human intent understanding, primarily due to the challenges of collecting real-world public discussion data and constructing a robust evaluation pipeline. To bridge this gap, we introduce \bench, the first dynamic, live evaluation benchmark specifically designed for intent understanding, particularly in the consumer domain. \bench is the largest and most diverse benchmark of its kind, supporting real-time updates while preventing data contamination through an automated curation pipeline.

意图理解大模型评测消费者行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。