揭示大模型单轮与多轮对话能力的差距,提出新评测与训练方案。
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
- 构建可对比的多轮对话评测基准TurnWiseEval。
- 仅用1万条多轮数据训练,性能提升12%。
- 适合关注对话系统长期交互能力的研究者。
多轮对话是语言模型交互的重要方式,但现有训练与评估数据主要聚焦单轮场景,未能捕捉长对话的额外复杂性。为探究单轮与多轮能力的差距,我们提出了新的评测基准TurnWiseEval,可直接与单轮对话评估对比。通过成对比较,隔离出多轮对话特有能力。同时,我们设计了合成多轮数据生成管道TurnWiseData,支持大规模多轮训练数据生成。在Olmo 3上的实验表明,使用多轮数据训练对实现强多轮对话性能至关重要;仅在后训练阶段加入10,000条多轮对话,即可在TurnWiseEval上带来12%的性能提升。
原文摘要 · Abstract (English)
Multi-turn conversations are a common and critical mode of language model interaction. However, current open training and evaluation data focus on single-turn settings, failing to capture the additional dimension of these longer interactions. To understand this multi-/single-turn gap, we first introduce a new benchmark, TurnWiseEval, for multi-turn capabilities that is directly comparable to single-turn chat evaluation. Our evaluation isolates multi-turn specific conversational ability through pairwise comparison to equivalent single-turn settings. We additionally introduce our synthetic multi-turn data pipeline TurnWiseData which allows the scalable generation of multi-turn training data. Our experiments with Olmo 3 show that training with multi-turn data is vital to achieving strong multi-turn chat performance, and that including as little as 10k multi-turn conversations during post-training can lead to a 12% improvement on TurnWiseEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。