arXiv:2601.15809cs.CL2026-01

通过干预模型激活,提升多语言摘要评估指标的准确性。

SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics

  • 在推理时引导模型激活向英语对齐,改善跨语言表现。
  • 多种指标在不同语言上相关性显著提升,最高增益达18.7%。
  • 适合关注多语言生成评估、模型可解释性的研究者。

越来越多的研究使用多语言语言模型进行自然语言生成任务,如摘要生成。该领域的主要实证瓶颈在于许多语言缺乏准确且稳健的评估指标,阻碍了进展。近期研究表明,多语言模型常以英语作为内部转换语言,与该中心语言的不一致会导致下游性能下降。受此启发,我们探讨这种偏差是否也存在于多语言神经评估指标中:能否通过将指标模型的激活引导至英语中心,提升其与人工判断的相关性?我们在编码器和解码器基线指标上测试了推理时干预方法,发现其在多种语言上均有效,显著提升了评估指标的表现。

原文摘要 · Abstract (English)

An increasing body of work has leveraged multilingual language models for Natural Language Generation tasks such as summarization. A major empirical bottleneck in this area is the shortage of accurate and robust evaluation metrics for many languages, which hinders progress. Recent studies suggest that multilingual language models often use English as an internal pivot language, and that misalignment with this pivot can lead to degraded downstream performance. Motivated by the hypothesis that this mismatch could also apply to multilingual neural metrics, we ask whether steering their activations toward an English pivot can improve correlation with human judgments. We experiment with encoder- and decoder-based metrics and find that test-time intervention methods are effective across the board, increasing metric effectiveness for diverse languages.

多语言评估神经指标推理干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。