arXiv:2605.26738cs.CL2026-05

让大模型学会社交语境中的对话技巧,提升实际交流效果。

KARMA: Karma-Aligned Reward Model Adaptation

论文配图:KARMA: Karma-Aligned Reward Model Adaptation
图 1 · 摘自论文原文
  • 用Reddit对话数据训练奖励模型,捕捉上下文影响的回应价值
  • 基于上下文的奖励信号使下游模型在语用任务上表现更优
  • 虽牺牲部分事实性,但显著改善对话自然度与社会适应性

人类交流依赖隐含的社会信号,有效性不仅取决于语义内容,更受语气、语境和对话规范影响。我们提出KARMA(Karma对齐奖励模型适配)框架,利用大规模社交互动数据训练语言模型掌握情境敏感的对话行为。KARMA在Reddit对话上训练奖励模型,预测响应价值时依赖上下文信息,并通过强化学习微调语言模型以提升语用相关任务表现。关键发现:表现最佳的奖励模型并非最准确预测Reddit点赞数的模型;完全依赖上下文的奖励模型虽对点赞预测能力较差,却带来显著更好的下游性能。我们在有无直接接触社交媒体数据的条件下评估了该方法,结果模型在语用行为上均有提升,且不良副作用明显减少。然而,无论是否接触原始数据,所有条件下事实性均下降,表明这一权衡源自奖励信号本身,而非训练数据噪声。

原文摘要 · Abstract (English)

Human communication depends on implicit social signals where effectiveness is shaped by tone, context, and conversational norms rather than semantic content alone. We introduce KARMA (Karma-Aligned Reward Model Adaptation), a framework for LLM learning of context-sensitive conversational behavior from large-scale social interaction data. KARMA trains a reward model on Reddit conversations to predict response valuation conditioned on context, and uses this signal to fine-tune language models via reinforcement learning to improve performance on pragmatics-mediated tasks. Critically, we find that the highest performing reward model does not lead to better downstream model alignment: a reward model relying exclusively on conversational context was a worse predictor of Reddit karma but yielded substantially better downstream performance. We evaluate the effects of KARMA applied to a downstream model with and without direct exposure to the social media data. The resulting models show improved pragmatics-mediated behaviors with largely mitigated undesirable side effects. Factuality is consistently diminished by KARMA across all conditions, including when the downstream model has no direct exposure to Reddit data, suggesting that this tension is embedded in the reward signal itself rather than introduced by noisy training data.

对话建模奖励模型语用学大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。