arXiv:2601.00263cs.CLcs.AI2026-01ACL被引 3

研究大模型生成多语言反事实样本的有效性与问题

Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation

  • 对比直接生成与翻译生成的多语言反事实样本质量
  • 发现跨语言编辑模式相似,但生成错误普遍存在
  • 多语言数据增强比跨语言增强更有效,尤其对低资源语言

反事实指对输入进行最小修改后导致模型预测改变的样本,是解释模型行为的有力方法。大语言模型在生成英文反事实样本方面表现优异,并具备多语言能力,但其在多语言场景下的有效性尚不明确。为此,我们对多语言反事实进行了全面研究:首先,在六种语言上评估直接生成与通过英文翻译生成的反事实样本。尽管翻译生成的样本有效性更高,但需更多修改,且质量仍不及原生英文样本。其次,发现高资源欧洲语言的编辑模式高度相似,表明跨语言扰动遵循共同策略。第三,识别并分类出四种在各类语言中反复出现的生成错误类型。最后,揭示多语言反事实数据增强(CDA)带来的模型性能提升大于跨语言CDA,尤其对低资源语言效果显著;然而,生成样本的缺陷限制了模型性能与鲁棒性的进一步提升。

原文摘要 · Abstract (English)

Counterfactuals refer to minimally edited inputs that cause a model's prediction to change, serving as a promising approach to explaining the model's behavior. Large language models (LLMs) excel at generating English counterfactuals and demonstrate multilingual proficiency. However, their effectiveness in generating multilingual counterfactuals remains unclear. To this end, we conduct a comprehensive study on multilingual counterfactuals. We first conduct automatic evaluations on both directly generated counterfactuals in the target languages and those derived via English translation across six languages. Although translation-based counterfactuals offer higher validity than their directly generated counterparts, they demand substantially more modifications and still fall short of matching the quality of the original English counterfactuals. Second, we find the patterns of edits applied to high-resource European-language counterfactuals to be remarkably similar, suggesting that cross-lingual perturbations follow common strategic principles. Third, we identify and categorize four main types of errors that consistently appear in the generated counterfactuals across languages. Finally, we reveal that multilingual counterfactual data augmentation (CDA) yields larger model performance improvements than cross-lingual CDA, especially for lower-resource languages. Yet, the imperfections of the generated counterfactuals limit gains in model performance and robustness.

反事实生成多语言模型模型解释数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。