用词元扭曲让多模态大模型更准确理解近视角下的场景
Token Warping Helps MLLMs Look from Nearby Viewpoints
- 用词元级反向扭曲替代像素扭曲,提升视角变换稳定性
- 在新提出的ViewBench上优于所有基线方法,包括生成式扭曲
- 适合研究多模态视觉推理与视角不变性的研究人员
多模态大语言模型(MLLM)在视觉推理中表现良好,但对视角变化仍敏感。传统像素级扭曲易受微小深度误差影响并引入几何失真。受人类心理意象理论启发,本文探索基于ViT的MLLM中的图像词元是否可作为视角变换的有效基础。对比前向与反向扭曲发现,反向词元扭曲通过在目标视角定义密集网格,并为每个网格点检索对应源视角词元,能显著提升稳定性并保持语义连贯性。在新提出的ViewBench基准测试中,词元级扭曲使MLLM在近视角推理中表现更可靠,持续优于所有基线,包括像素级扭曲、空间微调模型及生成式扭曲方法。
原文摘要 · Abstract (English)
Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as pixel-wise warping is highly sensitive to small depth errors and often introduces geometric distortions. Drawing on theories of mental imagery that posit part-level structural representations as the basis for human perspective transformation, we examine whether image tokens in ViT-based MLLMs serve as an effective substrate for viewpoint changes. We compare forward and backward warping, finding that backward token warping, which defines a dense grid on the target view and retrieves a corresponding source-view token for each grid point, achieves greater stability and better preserves semantic coherence under viewpoint shifts. Experiments on our proposed ViewBench benchmark demonstrate that token-level warping enables MLLMs to reason reliably from nearby viewpoints, consistently outperforming all baselines including pixel-wise warping approaches, spatially fine-tuned MLLMs, and a generative warping method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。