arXiv:2605.19274cs.CL2026-05

多语言模型用英文解释非英文输入时,解释看似流畅却不够真实。

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

  • 用英文作为解释枢纽,让模型找到的关键词与人类判断更接近
  • 但这些关键词对模型决策的因果支撑力下降,最多弱化5.7倍
  • 适合关注解释真实性而非表面匹配的研究者参考

在多语言部署中,常以英语解释来审计非英语输入。本文评估提取式解释(即模型识别输入词段并生成理由)发现:英语枢纽解释虽与人工理由的词段重合度更高,但其证据与模型预测的因果关联更弱,表现为全面性与充分性下降。在3个任务、5种语言及2类多语言大模型上,英语解释虽流畅但锚定松散,全面性最差时下降达5.7倍,而任务准确率保持稳定。对于社会语境分类,英语解释还丢失了关键语用线索,导致忠实度与词段一致性双降。建议在输入语言中审计解释,报告多维度忠实度指标,并将英语解释视为沟通摘要而非决策溯源。

原文摘要 · Abstract (English)

LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationale'' and uncover a systematic trade-off: English-pivot explanations can achieve higher span agreement with human rationales while their evidence becomes less causally grounded in the model's prediction, as measured by both comprehensiveness and sufficiency. Across 3 tasks, 5~languages, and 2~multilingual LLM families, we find that English explanations frequently produce fluent but loosely anchored rationales, with comprehensiveness degrading by up to 5.7x relative to native-language conditions - even as task accuracy remains stable across settings. For socially nuanced classification, English pivots also fail to preserve pragmatic cues, reducing both faithfulness and span agreement. We recommend auditing explanations in the input language, reporting multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than faithful decision traces.

可解释性多语言忠实度大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。