用偏好优化提升多语言反事实解释的准确性与简洁性
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

- 通过偏好优化将解释的准确性和简洁性转化为可学习信号
- 在7种语言上平均提升有效性12.55%,且不牺牲简洁性
- 适合关注多语言模型可解释性的研究者与应用开发者
自生成反事实解释(SCEs)是大语言模型(LLMs)生成的最小修改输入,能翻转自身预测,提供对黑箱模型行为的因果解释。然而,将其扩展到非英语语言仍具挑战:现有方法在非主导语言中难以生成有效解释,且有效性和简洁性之间存在持续权衡。我们提出Macro,一种基于直接偏好优化(DPO)的多语言SCE生成框架,采用复合评分函数构建偏好对,将该权衡转化为可度量的偏好信号。在四种LLM和七种语系多样的语言上的实验表明,Macro相比思维链基线平均提升有效性12.55%,且不降低简洁性,同时避免了翻译基线的严重简洁性破坏。相较于监督微调,Macro在两个指标上均表现更优,证实显式偏好优化对平衡该权衡至关重要。进一步分析显示,Macro增强了跨语言扰动对齐,并缓解常见生成错误。结果表明,偏好优化是提升多语言模型解释能力的有前景方向。
原文摘要 · Abstract (English)
Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior. Yet extending them beyond English remains challenging: existing methods struggle to produce valid SCEs in non-dominant languages, and a persistent trade-off between validity and minimality undermines explanation quality. We introduce Macro, a preference alignment framework that applies Direct Preference Optimization (DPO) to multilingual SCE generation, using a composite scoring function to construct preference pairs that effectively translate the trade-off into measurable preference signals. Experiments across four LLMs and seven typologically diverse languages show that Macro improves validity by 12.55\% on average over the chain-of-thought baseline without degrading minimality, while avoiding the severe minimality violations of the translation-based baseline. Compared to supervised fine-tuning, Macro achieves superior performance on both metrics, confirming that explicit preference optimization is essential for balancing this trade-off. Further analyses reveal that Macro increases cross-lingual perturbation alignment and mitigates common generation errors. Our results highlight preference optimization as a promising direction for enhancing multilingual model explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。