arXiv:2512.24693cs.CL2025-12

通过多轮对比增强训练,提升对话质量评估模型性能。

MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models

  • 设计多步指令对比方法,生成跨多轮的对话对比数据
  • 在多轮对话评估中优于基线,且保持单轮评估性能
  • 适合需要高质量多轮对话评估的研究者使用

评估多轮对话质量对构建高性能大语言模型至关重要,但传统方法依赖昂贵的人工评价。多轮奖励模型(RMs)提供了可扩展的替代方案,但在自动评估方面仍落后。我们发现,现有偏好数据集仅基于最后一轮对比响应,难以捕捉多轮交互的细微差别。为此,提出无监督的数据增强策略MUSIC,通过合成跨越多个回合的对比对话对,增强训练信号。在Skywork偏好数据集上,基于Gemma-2-9B-Instruct训练多轮奖励模型,实验证明MUSIC增强的模型在多轮对话评估中更贴近先进专有大模型评委的判断,同时未牺牲单轮基准测试表现。

原文摘要 · Abstract (English)

Evaluating the quality of multi-turn conversations is crucial for developing capable Large Language Models (LLMs), yet remains a significant challenge, often requiring costly human evaluation. Multi-turn reward models (RMs) offer a scalable alternative and can provide valuable signals for guiding LLM training. While recent work has advanced multi-turn \textit{training} techniques, effective automated \textit{evaluation} specifically for multi-turn interactions lags behind. We observe that standard preference datasets, typically contrasting responses based only on the final conversational turn, provide insufficient signal to capture the nuances of multi-turn interactions. Instead, we find that incorporating contrasts spanning \textit{multiple} turns is critical for building robust multi-turn RMs. Motivated by this finding, we propose \textbf{MU}lti-\textbf{S}tep \textbf{I}nstruction \textbf{C}ontrast (MUSIC), an unsupervised data augmentation strategy that synthesizes contrastive conversation pairs exhibiting differences across multiple turns. Leveraging MUSIC on the Skywork preference dataset, we train a multi-turn RM based on the Gemma-2-9B-Instruct model. Empirical results demonstrate that our MUSIC-augmented RM outperforms baseline methods, achieving higher alignment with judgments from advanced proprietary LLM judges on multi-turn conversations, crucially, without compromising performance on standard single-turn RM benchmarks.

奖励模型多轮对话数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。