arXiv:2505.20451cs.CL2025-05EMNLP被引 2

用对话行为和交际原则提升大模型对多轮对话的评判准确率

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries

  • 基于对话行为与交际准则构建评判框架
  • 75%的对话偏好可由行为或准则差异区分
  • 适合需要精准评估多轮对话质量的研究者

当前大语言模型广泛用作其他模型输出的评判者。然而,真实人机对话通常多轮、主题多样、意图多变。本文提出Amulet框架,利用对话行为(dialog acts)与交际准则(maxims)提升大模型在复杂多轮对话中的评判准确性。该框架分析对话中沟通结构与意图变化,并评估回复是否满足交际原则。在四个挑战性数据集上,实验发现人类在60%至70%的对话轮次中会改变意图,且75%的偏好响应可通过对话行为或交际准则加以区分,验证了其判别价值。Amulet既可独立作为单模型裁判,也可集成为多模型评审团,在所有四组数据上均显著优于基线方法。

原文摘要 · Abstract (English)

Today, large language models are widely used as judges to evaluate responses from other language models. Hence, it is imperative to benchmark and improve these LLM-judges on real-world language model usage: a typical human-assistant conversation is lengthy, and shows significant diversity in topics, intents, and requirements across turns, e.g. social interactions, task requests, feedback. We present Amulet, a framework that leverages pertinent linguistic concepts of dialog-acts and maxims to improve the accuracy of LLM-judges on preference data with complex, multi-turn conversational context. Amulet presents valuable insights about (a) the communicative structures and intents present in the conversation (dialog acts), and (b) the satisfaction of conversational principles (maxims) by the preference responses, and uses them to make judgments. On four challenging datasets, Amulet shows that (a) humans frequently (60 to 70 percent of the time) change their intents from one turn of the conversation to the next, and (b) in 75 percent of instances, the preference responses can be differentiated via dialog acts and/or maxims, reiterating the latter's significance in judging such data. Amulet can be used either as a judge by applying the framework to a single LLM, or integrated into a jury with different LLM judges; our judges and juries show strong improvements on relevant baselines for all four datasets.

对话评估大模型裁判多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。