研究多视频摘要中位置偏差,发现模型输出受视频输入顺序影响。
A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs

- 构建包含四类场景的多视频基准数据集,测试不同模型输入位置影响。
- 发现中间视频摘要质量普遍偏低,且偏差随模型和领域变化。
- 提示词优化可缓解部分偏差,但无法彻底消除位置敏感性。
多模态大语言模型(MLLMs)在视频理解中应用日益广泛,但其在多视频输入下的可靠性仍不清晰。本文研究多视频摘要中的位置偏差问题,即相同内容因输入位置不同导致摘要质量变化。基于ActivityNet与新闻视频构建基准数据集,涵盖烹饪、居家、休闲、新闻四类场景,支持双视频与四视频输入。评估九个开源与专有MLLMs,采用覆盖度、方向性位置偏差(DPB)与中-边差距(MEG)三个互补指标量化位置效应。结果表明:位置影响具有领域与模型依赖性,方向性偏差虽小,但中间位置表现普遍较差;增加视觉或生成资源并未均匀缓解不平衡。进一步分析提示词级缓解方法,发现仍不足以实现完全鲁棒性。结论指出,当前多视频摘要对输入顺序仍敏感,亟需发展更鲁棒的顺序无关系统。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly used for video understanding, yet their reliability under multi-video inputs remains poorly understood. We study positional bias in multi-video summarization, where the quality of a per-video summary can change with the video's input slot even when the underlying content is unchanged. We construct a benchmark from ActivityNet and News videos, covering Cooking, Domestic, Leisure, and News settings with two- and four-video inputs. We evaluate nine open-source and proprietary MLLMs and measure position effects with three complementary metrics: Coverage, Directional Positional Bias (DPB), and Middle-Edge Gap (MEG). Our results show that positional effects are domain- and model-dependent: signed directional bias can be small even when middle positions underperform, and increasing visual or generation budget does not uniformly remove the imbalance. We further analyze prompt-level mitigation methods. Together, the results show that multi-video summarization remains sensitive to input protocol and position, motivating more robust order-invariant multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。