测试多语言大模型在多方对话中的泛化能力,发现现有模型表现不佳。
Can MLLMs Generalize to Multi-Party dialog? Exploring Multilingual Response Generation in Complex Scenarios
- 构建首个面向多方对话的多语言平行数据集XMP,含三名以上参与者
- 大模型在多方场景下泛化能力差,微调后仅70B模型提升1%性能
- 跨语言训练反而降低效果,多语言互补性不成立
当前多语言大模型仍集中于简单问答任务,难以应对复杂对话结构。本文提出新问题:多语言大模型能否泛化到复杂对话?监督微调能否恢复该能力?多语言互补性是否有效?为此,我们构建XMP——首个基于多方播客对话的高质量多语言平行数据集,多数样本包含三名以上参与者,讨论广泛话题。实验表明:(R1)多语言大模型在多方对话中无法泛化;(R2)在XMP上微调仅带来微弱改善,70B模型相比8B模型最多提升1%绝对指标;(R3)混合语言微调通常有害,任何收益均有限且仅在70B模型中偶发出现。
原文摘要 · Abstract (English)
Current multilingual large language models(MLLMs) still focus on simple question-answering formats, often overlooking more complex dialogue scenarios. In other words, their capabilities of multilingual large models have yet to be validated in dialogue tasks with intricate structures. We therefore ask, Q1: How well do LLMs generalize to more complex dialog scenarios? Q2: Can supervised fine-tuning on a high-quality parallel benchmark restore this ability? Q3: Does the "multilingual complementarity" effect survive in the setting? To answer these questions, we introduce XMP, a high-quality parallel Multilingual dataset sourced from Multi-party Podcast dialogues, which is the first parallel dataset focusing on multi-party dialogue scenarios. Most samples in the dataset feature three or more participants, discussing a wide range of topics. Through extensive experiments, we find that, R1: MLLMs fail to generalize to multi-party setting, R2 Fine-tuning on XMP improves only marginally, with the 70B model achieving at most a 1% absolute gain over its 8B counterpart; R3: Mixing languages during SFT is usually detrimental, with any benefits being marginal and limited to isolated cases in the 70B model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。