诊断多视角人-物理解能力,揭示模型依赖单图语义平均而非真实跨视融合。
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity

- 分单视充分与多视必要两类任务,精准测试跨视推理能力
- 17个主流模型在冗余视图下准确率下降超40%,暴露注意力干扰问题
- 适合研究多模态模型跨视角理解机制的学者使用
多模态大语言模型在单图像感知上表现优异,但在复杂跨视角人本场景中的推理能力尚未得到充分验证。现有基准采用固定“帧集合”评估,混淆了模型对视觉干扰的鲁棒性与真实跨视证据融合能力。为此,我们提出SSMNBench,一个包含3,300个精心设计问答对的诊断基准,用于跨视角人与人-物理解。该基准将任务分为单视充分(SVS)与多视必要(MVN)两类。通过对17个先进MLLMs系统性扰动视图可用性,发现模型在冗余视图下出现严重“干扰退化”(SVS),且无法整合多摄像头间的碎片化几何信息(MVN)。评估表明,当前模型依赖多个单图语义平均和视图偏好,而非真正的跨视合成。该工作为未来跨视感知多模态架构的发展提供了严谨诊断框架。代码已开源。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view human-centric scenes remains largely unverified. Current multi-view benchmarks evaluate models using a fixed "bag of frames" and thus conflate a model's robustness to visual distraction with its genuine ability to fuse fragmented cross-view evidence. To address this issue, we introduce SSMNBench, a diagnostic benchmark comprising 3,300 curated QA pairs for cross-view human and human-object understanding. SSMNBench uniquely categorizes tasks into Single-View Sufficiency (SVS) and Multi-View Necessity (MVN). By systematically perturbing view availability across 17 state-of-the-art MLLMs, critical limitations are revealed: models suffer from severe "distraction degradation" when presented with redundant views (SVS), and fail to integrate fragmented geometric evidence across cameras (MVN). Our evaluations demonstrate that modern MLLMs rely on multiple single-image semantic averaging and view preference rather than genuine cross-view synthesis. By exposing these fundamental vulnerabilities, SSMNBench provides a rigorous diagnostic framework to drive the advancement of future cross-view-aware multimodal architectures. The code is available at: $ \href{https://github.com/gtc-gh/SSMNBench}{\text{SSMNBench}} $
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。