让大模型学会像人一样用对话形成沟通惯例,提升交流效率。
Post-training for Efficient Communication via Convention Formation
- 通过针对性微调人类对话中的惯例形成案例,训练模型适应多轮交互。
- 在两个新设计的评估任务中,微调后模型的惯例形成能力显著提升。
- 适合研究语言演化、对话系统优化的研究者和工程师参考。
人类在多轮对话中会不断调整语言,自发形成临时沟通惯例,从而提高交流效率。相比之下,已有研究表明大语言模型(LLM)不具备这种自然行为。本文提出一种后训练方法,通过在人工识别的惯例形成示范数据上进行针对性微调,使模型具备该能力。我们设计了两个新基准来评估这一能力:一是基于认知动机的聚焦性交互基准,能稳定诱发人类强烈的惯例形成趋势;二是反映真实场景的文档引导型指代补全任务,体现自然状态下的惯例形成行为。实验结果表明,经过后训练的大模型在这两项任务上的惯例形成能力均有显著提升。
原文摘要 · Abstract (English)
Humans communicate with increasing efficiency in multi-turn interactions, by adapting their language and forming ad-hoc conventions. In contrast, prior work shows that LLMs do not naturally show this behavior. We develop a post-training process to develop this ability through targeted fine-tuning on heuristically identified demonstrations of convention formation. We evaluate with two new benchmarks focused on this capability. First, we design a focused, cognitively-motivated interaction benchmark that consistently elicits strong convention formation trends in humans. Second, we create a new document-grounded reference completion task that reflects in-the-wild convention formation behavior. Our studies show significantly improved convention formation abilities in post-trained LLMs across the two evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。