arXiv:2510.20358cs.CL2025-10被引 4

仅用对话数据训练小模型,难以实现真正沟通能力。

Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)

  • 基于对话数据预训练,再用不同策略微调以提升沟通性。
  • 在最小配对对话延续任务中表现优异,但标准测试指标落后。
  • 直接偏好优化(DPO)效果优于强化学习,适合对话生成研究者。

我们研究了仅在对话数据上预训练是否能使小型语言模型具备形式和功能上的沟通能力。基于此预训练的 llamalogue 模型,我们采用多种微调策略,旨在使模型生成更具沟通性的文本。尽管模型在多数标准 BabyLM 基准测试中表现不佳,但在最小配对设置下的对话延续预测任务中表现出色。虽然 PPO 微调对模型效果有混合甚至负面作用,但直接偏好优化(DPO)进一步提升了其在自定义对话基准上的表现。

原文摘要 · Abstract (English)

We investigate whether pre-training exclusively on dialogue data results in formally and functionally apt small language models. Based on this pre-trained llamalogue model, we employ a variety of fine-tuning strategies to enforce "more communicative" text generations by our models. Although our models underperform on most standard BabyLM benchmarks, they excel at dialogue continuation prediction in a minimal pair setting. While PPO fine-tuning has mixed to adversarial effects on our models, DPO fine-tuning further improves their performance on our custom dialogue benchmark.

BabyLM对话生成模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。