arXiv:2505.05970cs.CL2025-05被引 4

用对话成功作为信号训练语言模型,模拟儿童学语言过程。

Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models

  • 以对话成功为奖励,引导模型在纯语言场景中优化表达。
  • 模型行为变化符合认知约束,但语言评估未显著提升。
  • 适合研究语言习得机制的计算模型与交互学习方向。

我们提出一种受儿童语言习得启发的交互式语言模型训练方法。在单轮对话中,说话者尝试向听者传递信息,若沟通成功则获得奖励。与以往基于图文配对的参考游戏不同,本工作在纯语言问答场景中定义沟通成功。首先,可行性研究显示该奖励能间接反映语法正确性。其次,通过强化学习微调语言模型,发现通信通道的认知合理限制可引发可解释的说话行为变化。然而,当前训练方式尚未在语言评估上带来改善。论文进一步建议任务设计与训练配置的调整,以在未来研究中更有效验证交互对语言学习的促进作用。

原文摘要 · Abstract (English)

We propose a method for training language models in an interactive setting inspired by child language acquisition. In our setting, a speaker attempts to communicate some information to a listener in a single-turn dialogue and receives a reward if communicative success is achieved. Unlike earlier related work using image--caption data for interactive reference games, we operationalize communicative success in a more abstract language-only question--answering setting. First, we present a feasibility study demonstrating that our reward provides an indirect signal about grammaticality. Second, we conduct experiments using reinforcement learning to fine-tune language models. We observe that cognitively plausible constraints on the communication channel lead to interpretable changes in speaker behavior. However, we do not yet see improvements on linguistic evaluations from our training regime. We outline potential modifications to the task design and training configuration that could better position future work to use our methodology to observe the benefits of interaction on language learning in computational cognitive models.

语言模型交互学习奖励设计认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。