arXiv:2506.11063cs.CLcs.AI2025-06EMNLP被引 11

发现多模态检索生成中证据位置影响模型表现,提出量化方法揭示偏见机制。

Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation

  • 通过控制实验发现证据位置呈U形影响准确率
  • 提出PSI_p指数,量化位置敏感度并揭示注意力分配模式
  • 适合关注多模态系统可靠性与公平性的研究者参考

多模态检索增强生成(RAG)系统在知识密集型和开放域任务中日益重要。随着检索复杂度上升,确保系统鲁棒性至关重要。然而,现有RAG模型对证据呈现顺序高度敏感,导致性能不稳定和推理偏见,尤其当检索项数量或模态多样性增加时更为明显。本文首次全面研究多模态RAG中的位置偏见问题。通过在纯文本、纯图像及混合模态任务上的受控实验,观察到准确率随证据位置呈现一致的U形曲线。为量化该偏见,提出位置敏感度指数(PSI_p),并构建可视化框架追踪解码器各层注意力分布。结果表明,多模态交互比单模态显著加剧位置偏见,且偏见随检索范围呈对数增长。这些发现为RAG的位置感知分析提供理论与实证基础,强调需引入证据重排序或去偏策略以构建更可靠、公平的生成系统。

原文摘要 · Abstract (English)

Multimodal Retrieval-Augmented Generation (RAG) systems have become essential in knowledge-intensive and open-domain tasks. As retrieval complexity increases, ensuring the robustness of these systems is critical. However, current RAG models are highly sensitive to the order in which evidence is presented, often resulting in unstable performance and biased reasoning, particularly as the number of retrieved items or modality diversity grows. This raises a central question: How does the position of retrieved evidence affect multimodal RAG performance? To answer this, we present the first comprehensive study of position bias in multimodal RAG systems. Through controlled experiments across text-only, image-only, and mixed-modality tasks, we observe a consistent U-shaped accuracy curve with respect to evidence position. To quantify this bias, we introduce the Position Sensitivity Index ($PSI_p$) and develop a visualization framework to trace attention allocation patterns across decoder layers. Our results reveal that multimodal interactions intensify position bias compared to unimodal settings, and that this bias increases logarithmically with retrieval range. These findings offer both theoretical and empirical foundations for position-aware analysis in RAG, highlighting the need for evidence reordering or debiasing strategies to build more reliable and equitable generation systems.

多模态RAG偏见分析注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。