模型每一步推理是否紧扣视觉输入,决定了其在新场景下的泛化能力。
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
- 通过衡量模型每一步推理与视觉变化的一致性,定义了时间对齐度指标。
- 对齐度越高,模型在分布外数据上的表现越强,相关性达0.83(p=0.003)。
- 即使参数相同,不同模型的对齐度差异可达10.8个百分点,是独立能力维度。
我们发现长时程视觉语言模型存在一种行为规律:保持时间上对齐信念的模型泛化能力更强。标准基准仅评估最终答案准确率,掩盖了模型如何使用视觉信息;一个模型可能猜对答案,但其逐步推理完全脱离视觉输入。我们将其形式化为长时程行为忠实度,一种可实证测量的属性,用于量化模型中间推理是否与动态视觉状态保持一致。在三个长时程基准上的八种模型中,我们证明时间对齐质量是鲁棒性的关键预测因子:步骤对齐率(SGR)与分布外保留率的相关性为 $r = 0.83$(置换检验 $p = 0.003$),该关系在容量匹配模型中依然成立,且无法由规模或分布内准确率解释。关键的是,在参数量相同的7B模型中,对齐度差异高达10.8个百分点,表明其为独立的能力轴。多重稳健性检验确认该信号反映真实的视觉依赖:反事实轨迹使SGR下降26–41个百分点,跨架构验证器一致性达 $ρ=0.96$,随机推理得分接近随机水平(∼18%),且即使无显式推理披露,预测力仍强($r = 0.78$)。
原文摘要 · Abstract (English)
We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standard benchmarks measure only final-answer accuracy, which obscures how models use visual information; a model can guess correctly while its step-by-step reasoning is entirely unanchored to the visual input. We formalize this as behavioral faithfulness over long horizons, an empirically measurable property that quantifies whether a model's intermediate reasoning remains consistent with the evolving visual state. Across eight models on three long-horizon benchmarks, we demonstrate that temporal grounding quality is a leading indicator of robustness: the Step Grounding Rate (SGR) predicts out-of-distribution retention with $r = 0.83$ (permutation test $p = 0.003$), a relationship that holds within capacity-matched models and cannot be explained by scale or in-distribution accuracy. Critically, grounding quality varies by up to 10.8 percentage points within parameter-matched 7B models despite similar accuracy, revealing it as an independent axis of model capability. Multiple robustness checks confirm the signal reflects genuine visual reliance: counterfactual traces drop SGR by 26--41 percentage points, cross-architecture verifiers agree at $ρ= 0.96$, random reasoning scores near chance ($\sim 18\%$), and the predictor remains strong even without explicit reasoning disclosure ($r = 0.78$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。