arXiv:2508.06482cs.CLcs.AI2025-08中稿 · COLM被引 6

让大模型学会像人一样用对话形成沟通惯例,提升交流效率。

Post-training for Efficient Communication via Convention Formation

  • 通过针对性微调人类对话中的惯例形成案例,训练模型适应多轮交互。
  • 在两个新设计的评估任务中,微调后模型的惯例形成能力显著提升。
  • 适合研究语言演化、对话系统优化的研究者和工程师参考。

人类在多轮对话中会不断调整语言,自发形成临时沟通惯例,从而提高交流效率。相比之下,已有研究表明大语言模型(LLM)不具备这种自然行为。本文提出一种后训练方法,通过在人工识别的惯例形成示范数据上进行针对性微调,使模型具备该能力。我们设计了两个新基准来评估这一能力:一是基于认知动机的聚焦性交互基准,能稳定诱发人类强烈的惯例形成趋势;二是反映真实场景的文档引导型指代补全任务,体现自然状态下的惯例形成行为。实验结果表明,经过后训练的大模型在这两项任务上的惯例形成能力均有显著提升。

原文摘要 · Abstract (English)

Humans communicate with increasing efficiency in multi-turn interactions, by adapting their language and forming ad-hoc conventions. In contrast, prior work shows that LLMs do not naturally show this behavior. We develop a post-training process to develop this ability through targeted fine-tuning on heuristically identified demonstrations of convention formation. We evaluate with two new benchmarks focused on this capability. First, we design a focused, cognitively-motivated interaction benchmark that consistently elicits strong convention formation trends in humans. Second, we create a new document-grounded reference completion task that reflects in-the-wild convention formation behavior. Our studies show significantly improved convention formation abilities in post-trained LLMs across the two evaluation methods.

语言演化对话系统后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。