arXiv:2512.02665cs.CL2025-12

输入顺序影响大模型摘要的语义一致性,首篇文档权重最高。

Input Order Shapes LLM Semantic Alignment in Multi-Document Summarization

  • 通过六种输入顺序对比,发现模型更依赖第一篇文档的语义。
  • 首篇文档生成的摘要与原文语义相似度高出23.7%,显著优于后两篇。
  • 适用于关注提示工程偏差、信息公平性的研究者与应用开发者。

大型语言模型(LLMs)现被用于谷歌AI概览等多文档摘要场景,但其对输入文档的加权方式尚不明确。以堕胎相关新闻为例,构建40组‘支持-中立-反对’文章三元组,将每组按六种顺序排列,使用Gemini 2.5 Flash生成中立概述。采用ROUGE-L(词汇重叠)、BERTScore(语义相似度)和SummaC(事实一致性)评估摘要质量。单因素方差分析显示,所有立场下BERTScore均呈现显著首因效应,说明摘要与首篇文档语义更一致。成对比较进一步表明,位置1显著区别于位置2和3,而位置2与3无差异,证实模型存在对首篇文档的选择性偏好。该结果揭示了依赖大模型生成概览的应用及智能体系统中的潜在风险。

原文摘要 · Abstract (English)

Large language models (LLMs) are now used in settings such as Google's AI Overviews, where it summarizes multiple long documents. However, it remains unclear whether they weight all inputs equally. Focusing on abortion-related news, we construct 40 pro-neutral-con article triplets, permute each triplet into six input orders, and prompt Gemini 2.5 Flash to generate a neutral overview. We evaluate each summary against its source articles using ROUGE-L (lexical overlap), BERTScore (semantic similarity), and SummaC (factual consistency). One-way ANOVA reveals a significant primacy effect for BERTScore across all stances, indicating that summaries are more semantically aligned with the first-seen article. Pairwise comparisons further show that Position 1 differs significantly from Positions 2 and 3, while the latter two do not differ from each other, confirming a selective preference for the first document. The findings present risks for applications that rely on LLM-generated overviews and for agentic AI systems, where the steps involving LLMs can disproportionately influence downstream actions.

大模型摘要输入顺序语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。