提升大模型多轮对话的一致性,让其回答更稳定可靠。
Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions
- 引入位置加权一致性度量,关注早期回复的稳定性与后续修复能力。
- 构建多领域挑战性数据集,全面评估模型在复杂追问下的表现。
- 通过融合内部置信度生成新方法,显著增强回答稳定性。
大语言模型在各类任务中表现出色,但在高风险场景中部署时,需确保多轮交互中的行为一致性和连贯性。本文提出一个全面的评估与改进框架,包含三项关键贡献:首先,提出位置加权一致性(PWC)度量,用于捕捉多轮交互中早期阶段的稳定性与后续恢复模式;其次,构建了涵盖多样领域和难度级别的MT-Consistency基准数据集,专门用于评估模型在各种挑战性追问情境下的表现;第三,提出置信度感知生成(CARG)框架,通过在生成过程中显式整合模型内部置信度得分,显著提升响应稳定性。实验结果表明,CARG在不牺牲准确率的前提下大幅提升了响应稳定性,为大模型在关键真实场景中的可靠应用提供了可行路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities across various tasks, but their deployment in high-stake domains requires consistent and coherent behavior across multiple rounds of user interaction. This paper introduces a comprehensive framework for evaluating and improving LLM response consistency, making three key contributions. Code and data are available at: https://github.com/yubol-bobo/MT-Consistency. First, we introduce Position-Weighted Consistency (PWC), a metric designed to capture both the importance of early-stage stability and recovery patterns in multi-turn interactions. Second, we present MT-Consistency, a carefully curated benchmark dataset spanning diverse domains and difficulty levels, specifically designed to evaluate LLM consistency under various challenging follow-up scenarios. Third, we introduce Confidence-Aware Response Generation (CARG), a framework that significantly improves response stability by explicitly integrating internal model confidence scores during the generation process. Experimental results demonstrate that CARG significantly improves response stability without sacrificing accuracy, offering a practical path toward more dependable LLM behavior in critical, real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。