评测大模型在复杂体育对话中的多轮理解能力,发现闭源模型更优。
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
- 基于体育解说文本构建多轮对话基准,评估长程依赖与动机传递。
- 闭源模型显著优于开源模型,显式推理可提升复杂对话处理能力。
- 揭示注意力机制因特殊标记失效导致长对话性能下降的机理。
大型语言模型(LLMs)如ChatGPT已广泛应用于真实对话场景,但其在处理长而复杂的多轮对话时的鲁棒性仍受质疑,尤其在频繁动机转移和复杂跨轮依赖方面。现有基准未能充分反映这些缺陷。本文提出MARS-Bench——一个基于体育赛事逐句解说文本构建的多轮体育现实场景对话基准,专门用于评估三类关键能力:超长多轮、交互式多轮和跨轮任务。在MARS-Bench上的实验表明,闭源模型显著优于开源模型,显式推理能显著提升模型在复杂对话中的鲁棒性,且模型在处理动机转移与复杂跨轮依赖时仍面临重大挑战。此外,通过对Qwen2.5-7B-Instruction进行注意力可视化分析,揭示了特殊标记导致注意力衰减是性能下降的关键机制。
原文摘要 · Abstract (English)
Large Language Models (\textbf{LLMs}), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs' robustness, especially in handling long complex dialogue sessions, including frequent motivation transfer, sophisticated cross-turn dependency, is criticized all along. Nevertheless, no existing benchmarks can fully reflect these weaknesses. We present \textbf{MARS-Bench}, a \textbf{M}ulti-turn \textbf{A}thletic \textbf{R}eal-world \textbf{S}cenario Dialogue \textbf{Bench}mark, designed to remedy the gap. MARS-Bench is constructed from play-by-play text commentary so to feature realistic dialogues specifically designed to evaluate three critical aspects of multi-turn conversations: Ultra Multi-turn, Interactive Multi-turn, and Cross-turn Tasks. Extensive experiments on MARS-Bench also reveal that closed-source LLMs significantly outperform open-source alternatives, explicit reasoning significantly boosts LLMs' robustness on handling long complex dialogue sessions, and LLMs indeed face significant challenges when handling motivation transfer and sophisticated cross-turn dependency. Moreover, we provide mechanistic interpretability on how attention sinks due to special tokens lead to LLMs' performance degradation when handling long complex dialogue sessions based on attention visualization experiment in Qwen2.5-7B-Instruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。