arXiv:2603.14217cs.CL2026-03中稿 · LREC 2026, officia…被引 1

用认知语言学视角重评对话系统评估,发现现有方法忽略深层语境一致性。

Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective

  • 从认知与语言学角度重构对话评估框架,强调连贯性与共享理解
  • 发现主流指标如BLEU与人类/大模型判断差异显著,无法捕捉对话缺陷
  • 适合关注对话质量评估、个性化生成系统的研究者参考

在认知科学和语言学理论中,对话被视为由连贯性、一致性和共享理解维系的协作活动,而非独立话语的链式组合。然而,当前开放域个性化对话系统普遍依赖表面相似度指标(如BLEU、ROUGE、F1)作为主要评估手段,难以反映对话的深层质量。本文以一个典型的检索增强型个性化对话框架LAPDOG为例,通过人工与大模型双重评判,揭示现有评估方法存在对话历史污染、检索故事与个人设定矛盾、回复不连贯等问题。结果显示,人类与大模型评价高度一致,却明显偏离词法相似性指标,凸显构建基于认知基础的评估方法的必要性。本研究为更可靠地衡量检索增强型对话系统提供了符合自然人际交流原则的新路径。

原文摘要 · Abstract (English)

In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for open-domain and personalized dialogue use surface-level similarity metrics (e.g., BLEU, ROUGE, F1) as one of their main reporting measures, which fail to capture these deeper aspects of conversational quality. We re-examine a notable retrieval-augmented framework for personalized dialogue, LAPDOG, as a case study for evaluation methodology. Using both human and LLM-based judges, we identify limitations in current evaluation practices, including corrupted dialogue histories, contradictions between retrieved stories and persona, and incoherent response generation. Our results show that human and LLM judgments align closely but diverge from lexical similarity metrics, underscoring the need for cognitively grounded evaluation methods. Broadly, this work charts a path toward more reliable assessment frameworks for retrieval-augmented dialogue systems that better reflect the principles of natural human communication.

对话评估认知语言学检索增强大模型评判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。