测试多模态大模型在不同视角下的空间一致性,发现其看似准确实则不稳。
CVT-Bench: Probing Spatial-State Integrity through Counterfactual Viewpoint Transformations
- 设计跨视角对抗性测试框架,检验模型对空间状态的持续判断能力。
- 5个顶尖模型在复杂场景中普遍出现联合不可实现的状态,准确性下降超40%。
- 提示词长度或位置不影响结果,说明问题出在模型内在机制而非外部干扰。
多模态大语言模型在孤立空间任务上表现良好,但其预测在不同视角和竞争场景下是否保持一致仍不明确。本文将此行为特性定义为「空间状态完整性」,提出CVT-Bench——一个涵盖两个领域(CVT-Synthetic、CVT-Real)、四种上下文情境(孤立、无关填充、属性填充、竞争场景)、三种表示形式(图像、文本/边界框、场景图)及十种视角条件(从0°到360°每隔45°共九个方位角,加顶部视角)的因子诊断套件,目标视图被隐藏。在200个场景与11,833个关系查询中,评估了反事实准确率、循环一致性、归一化存活系数、空间可实现性及上下文特异性干扰。五个前沿MLLMs表现出快速持久性丧失,尽管局部准确率高,却频繁生成联合不可实现的空间状态,且在竞争上下文与真实场景复杂度下失败加剧。无关与属性填充对照表明,提示长度、位置或场景访问丢失均无法解释该现象。平均而言,文本/边界框提升准确率、持久性与可实现性,场景图带来互补增益,但均未能消除不稳定性。因此,孤立空间准确率严重夸大了模型鲁棒性,确立空间状态完整性作为独立评测目标。完整基准与代码库将公开发布。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) perform strongly on isolated spatial tasks, but whether their predictions remain persistent and mutually coherent across viewpoints and competing scenes is unclear. We formalize this behavioral property as spatial-state integrity and introduce CVT-Bench, a factorial diagnostic suite spanning two domains (CVT-Synthetic, CVT-Real), four context regimes (CVT-Isolated, irrelevant filler, attribute filler, CVT-Competing), three representations (Image, Text/BBox, Scene Graph), and ten viewpoint conditions (nine azimuths from $0^\circ$ to $360^\circ$ at $45^\circ$ increments, plus Top) with target views withheld. Across 200 scenes and 11,833 relational queries, we measure counterfactual accuracy, cycle consistency, the normalized survival coefficient, spatial realizability, and context-specific interference. Five state-of-the-art MLLMs exhibit rapid persistence loss and frequently produce jointly unrealizable states despite high local accuracy, with failures amplified by competing contexts and natural-scene complexity. Matched irrelevant and attribute filler controls show that prompt length, position, or loss of scene access alone are insufficient explanations. On average, Text/BBox improves accuracy, persistence, and realizability, while Scene Graph provides complementary gains; neither eliminates the instability. Thus, isolated spatial accuracy substantially overestimates robustness, establishing spatial-state integrity as a distinct evaluation target. The complete benchmark and codebase will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。