评测大模型理解集体意图能力,推动其从执行指令到分析社会共识的跨越。
COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding
- 构建动态实时更新的基准COIN-BENCH,聚焦消费者领域集体讨论
- 20个主流大模型在深度、广度、信息量和正确性上表现有限,难达深层推理
- 提出树状结构与检索增强验证,提升对复杂意见分歧的解析能力
理解人类意图是大语言模型(LLM)面临的一项高阶认知挑战,需对嘈杂、冲突且非线性的公共讨论进行复杂推理。尽管现有模型能良好执行个体指令,但对其从多源公众讨论中提炼共识、化解矛盾、推断潜在趋势的集体意图理解能力仍缺乏系统评估。为此,我们提出COIN-BENCH——一个面向消费领域的动态、真实世界、持续更新的基准测试。不同于传统以交易结果为导向的评测,COIN-BENCH将意图定义为从显性场景到深层因果推理的分层认知结构。通过结合规则方法与大模型作为裁判的评估框架,引入COIN-TREE实现分层认知建模,并采用检索增强验证(COIN-RAG)保障分析精度。对20个先进大模型在深度、广度、信息量和正确性四个维度的评估显示,当前模型仅能处理表层聚合,仍难以实现复杂意图合成所需的分析深度。COIN-BENCH为推动大模型从被动指令执行者转变为具备专家级分析能力的集体话语解读者树立了新标准。
原文摘要 · Abstract (English)
Understanding human intent is a high-level cognitive challenge for Large Language Models (LLMs), requiring sophisticated reasoning over noisy, conflicting, and non-linear discourse. While LLMs excel at following individual instructions, their ability to distill Collective Intent - the process of extracting consensus, resolving contradictions, and inferring latent trends from multi-source public discussions - remains largely unexplored. To bridge this gap, we introduce COIN-BENCH, a dynamic, real-world, live-updating benchmark specifically designed to evaluate LLMs on collective intent understanding within the consumer domain. Unlike traditional benchmarks that focus on transactional outcomes, COIN-BENCH operationalizes intent as a hierarchical cognitive structure, ranging from explicit scenarios to deep causal reasoning. We implement a robust evaluation pipeline that combines a rule-based method with an LLM-as-the-Judge approach. This framework incorporates COIN-TREE for hierarchical cognitive structuring and retrieval-augmented verification (COIN-RAG) to ensure expert-level precision in analyzing raw, collective human discussions. An extensive evaluation of 20 state-of-the-art LLMs across four dimensions - depth, breadth, informativeness, and correctness - reveals that while current models can handle surface-level aggregation, they still struggle with the analytical depth required for complex intent synthesis. COIN-BENCH establishes a new standard for advancing LLMs from passive instruction followers to expert-level analytical agents capable of deciphering the collective voice of the real world. See our project page on COIN-BENCH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。