拆解多语言大模型任务中的语言角色,发现响应语言是性能关键
Disentangling Language Roles in Multilingual LLM Task Execution

- 构建全交叉语言三元组测试集,精确控制指令、内容、响应语言组合
- 27种语言组合共2430个实例,揭示响应语言错配导致主要性能下降
- 模型表现差异不只看错多少词,不同任务失败模式各异,适合多语言研究者
多语言大模型在指令、源内容与目标回应语言不一致时应用日益广泛。现有评估基准虽扩展了多语言指令遵循能力测试,但很少在完全交叉设计下分离这三个语言角色。本文提出MTM-Bench,一个语言条件化任务执行的受控基准,每个实例由三元组$(L_{\text{instr}}, L_{\text{content}}, L_{\text{resp}})$定义。在英语、西班牙语和中文之间,MTM-Bench枚举全部27种组合,每种模型包含2,430个实例,涵盖语义反转、最终状态提取与语言纯度更新实现。我们评估了20个前沿及开源大模型,使用分解指标:语义正确性、目标语言一致性、约束满足度、污染比例与联合成功率,并通过针对性人工审核验证评分。全交叉设计显示,性能退化由语言在任务结构中的角色决定,而非仅由错配数量决定。响应语言角色是主要变化轴,单一响应槽错配即造成最大退化。对比仅响应错与全错情况表明,错配数量并非单调预测因子,模型排序随系统而异。不同任务类别通过不同路径失败,说明仅关注语义正确性无法可靠衡量多语言任务执行。
原文摘要 · Abstract (English)
Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expanded multilingual instruction-following evaluation, but they rarely isolate these three roles within a fully crossed design. We introduce MTM-Bench, a controlled benchmark for language-conditioned task execution in which each instance is defined by a triplet \((L_{\text{instr}}, L_{\text{content}}, L_{\text{resp}})\). Across English, Spanish, and Chinese, MTM-Bench enumerates all 27 triplets and contains 2{,}430 instances per model across semantic reversal, final-state extraction, and language purity with update realization. We evaluate 20 frontier and open-weight LLMs using decomposed metrics for semantic correctness, target-language adherence, constraint satisfaction, contamination ratio, and joint success, with scoring validated by a targeted human audit. The fully crossed design reveals that degradation is organized by the role a language occupies in the task structure, not merely by mismatch count. The response-language role is the dominant axis of variation, and a single response-slot mismatch accounts for most degradation. The response-only and full-mismatch comparison suggests that mismatch count is not a monotonic predictor of difficulty, with model-level ordering varying across systems. Task families fail through distinct channels, showing that semantic correctness alone does not capture reliable multilingual task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。