研究多语言摘要中大模型的行为机制并提升生成质量
Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

- 构建覆盖24种语言的多目标跨语言摘要基准
- 发现翻译与摘要在后期层共同演化,错误也集中于此
- 通过英文化隐藏状态引导,跨语言摘要效果持续提升
多目标跨语言文本摘要(MTXLS)旨在将源文档同时摘要为多种目标语言,随着用户内容消费语言多样化,其重要性日益凸显,但相关研究仍较少。为此,我们提出多目标跨语言要素感知(MEA)基准,涵盖24种目标语言。我们在不同LLM上对比端到端与流水线方法,发现MTXLS性能仍显著落后于英语单语摘要。为深入理解大模型在MTXLS中的行为,我们设计分层分析框架,揭示翻译与摘要行为在后期层共同出现,而非明确分离;大部分任务相关处理及错误均集中于这些层。基于此,我们提出推理时激活引导方法,利用英语摘要的隐藏表示指导多语言生成。实验表明,该方法在多种目标语言上一致提升摘要质量。
原文摘要 · Abstract (English)
Multi-target cross-lingual text summarization (MTXLS), which summarizes a source document into multiple target languages, is increasingly important as users consume content in diverse languages, but remains underexplored. To address this gap, we introduce multi-target cross-lingual element-aware (MEA), a new MTXLS benchmark covering 24 target languages. We benchmark end-to-end and pipeline approaches across various LLMs and show that MTXLS performance still substantially lags behind English monolingual summarization. To better understand MTXLS in LLMs, we propose a layer-wise analysis framework for investigating how LLMs internally perform MTXLS. Our analyses suggest that translation and summarization behaviors emerge jointly within later layers rather than as distinctly decomposed stages. Most task-relevant processing occurs within these layers, and errors also tend to arise at similar depths. Motivated by these findings, we introduce an inference-time activation steering method that leverages hidden representations from English summarization to guide MTXLS generation. Experiments show that our method consistently improves MTXLS quality across target languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。