arXiv:2511.08394cs.CLcs.AI2025-11被引 6

用对话的动态轨迹做奖励信号,能有效提升大模型协作能力。

Interaction Dynamics as a Reward Signal for LLMs

  • 从对话嵌入轨迹的几何特征提取奖励信号,不依赖文本内容。
  • 仅用动态信号的模型准确率达68.20%,接近全文本基线的70.04%。
  • 结合文本与动态信号可达到80.17%最高性能,适合对话系统优化者。

大型语言模型在多轮对话中的对齐通常依赖于文本内容生成的奖励信号。然而,这种做法忽略了另一种丰富且互补的信号来源:交互本身的动力学特性。本文提出一种名为TRACE(基于轨迹的代理协作评估奖励)的新奖励信号,其来源于对话嵌入轨迹的几何属性——我们称之为“对话几何”。核心发现是,仅基于这些结构信号训练的奖励模型,在成对准确性上达到68.20%,与分析完整对话文本的强大基线模型(70.04%)相当。此外,结合交互动态与文本分析的混合模型表现最优,达到80.17%,证明两者具有互补性。该研究为交互场景提供了有力证据:在成功协作中,沟通方式与所说内容一样重要,同时构建了一个隐私友好的新框架,不仅用于对齐代理,还可作为诊断工具识别驱动高效协作的独特互动模式。

原文摘要 · Abstract (English)

The alignment of Large Language Models (LLMs) for multi-turn conversations typically relies on reward signals derived from the content of the text. This approach, however, overlooks a rich, complementary source of signal: the dynamics of the interaction itself. This paper introduces TRACE (Trajectory-based Reward for Agent Collaboration Estimation), a novel reward signal derived from the geometric properties of a dialogue's embedding trajectory--a concept we term 'conversational geometry'. Our central finding is that a reward model trained only on these structural signals achieves a pairwise accuracy (68.20%) comparable to a powerful LLM baseline that analyzes the full transcript (70.04%). Furthermore, a hybrid model combining interaction dynamics with textual analysis achieves the highest performance (80.17%), demonstrating their complementary nature. This work provides strong evidence that for interactive settings, how an agent communicates is as powerful a predictor of success as what it says, offering a new, privacy-preserving framework that not only aligns agents but also serves as a diagnostic tool for understanding the distinct interaction patterns that drive successful collaboration.

对话建模奖励设计大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。