用大模型评判图像序列时间逻辑,发现其判断结果受帧位置影响严重。
Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

- 用诊断测试揭示大模型在判断图像顺序时存在首尾帧偏差
- 同一故事不同排序下评分差异高达40%,暴露结构缺陷
- 适合关注视频理解评估的科研人员与模型开发者
随着生成式多模态技术从静态图像转向复杂交错的视觉叙事,一个基础性瓶颈浮现:评估危机。人类能自然感知故事的时间与逻辑脉络,而自动化评估系统却对序列连续性‘视而不见’,难以区分连贯叙事与语义错乱或矛盾的序列。本文指出当前多模态评估范式存在关键结构性缺陷,认为依赖大视觉语言模型(LVLMs)作为裁判,因架构偏见而根本受限。分析显示显著性能分裂:模型在孤立打分时表现尚可,但在需进行成对时间顺序判别时出现灾难性崩溃。这不仅是数据稀缺问题,更是结构问题。通过一系列诊断探针,我们发现系统性位置不对称——特别是首因效应与近因效应,即模型对故事的判断受帧位置影响远超语义一致性。这些偏差可能源于因果掩码与旋转位置编码。研究提示,现有基于Transformer的裁判模型本质上不适合长序列视觉推理。因此,呼吁多媒体领域摒弃以快照为中心的指标,转而构建真正面向时间感知的评估范式,将视觉序列视为统一逻辑结构而非无序帧集合。
原文摘要 · Abstract (English)
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。