新基准测试揭示大模型对话中意图追踪的隐藏差距
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

- 采用用户模拟、目标模型、独立评判三者分离的设计
- 发现顶尖模型在客观意图追踪上差9倍,主观表现却相近
- 适合评估对话系统长期一致性与真实用户体验
评估大语言模型作为多轮对话伙伴的能力,需要超越单轮测试的维度:人物一致性、意图演变追踪、情感动态以及多轮目标完成。我们提出EYT-Bench,一个以人类为中心的基准,其评估协议基于解耦的三方设计:基于人物设定的用户模拟器、被评估的目标模型(负责意图感知与回复生成)、以及独立可配置的LLM评判集成。在3400场对话与17个目标模型的测试中,揭示了以往基准忽略的四个发现:(i) 当前最先进闭源与开源模型在主观维度上无统计差异,但在客观意图追踪上差距最高达9倍;(ii) 长上下文人物设定下,推理能力对客观追踪呈相变式提升,而主观评分基本平坦;(iii) 人物格式显著影响轨迹发散程度,FICR(最终意图完成率)在Nemotron-USA超过0.95,而在PersonaMem-v2中为0.53至0.88之间;(iv) 17个模型中有16个出现预热效应。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal completion across many turns. We introduce EYT-Bench, a human-centered benchmark whose evaluation protocol is built around a decoupled three-party design: a persona-grounded user simulator, a target model evaluated on both intent perception and response generation, and an independent, configurable ensemble of LLM judges. Across 3,400 dialogues with 17 target models, EYT-Bench reveals four findings that previous benchmarks miss: (i) state-of-the-art closed and open-source models are statistically indistinguishable on subjective dimensions, but separate by up to 9x on objective intent-tracking; (ii) reasoning is a phase transition for objective tracking on long-context personas but is essentially flat on subjective scores; (iii) persona format strongly affects trajectory spread, FICR (final-intent completion rate) saturates above 0.95 on Nemotron-USA but ranges from 0.53 to 0.88 on PersonaMem-v2; and (iv) the warm-up effect is observed in 16 of 17 models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。