用深度图增强视觉模型,提升长时第一视角视频的空间推理能力。
Spatial-Conditioned Reasoning in Long-Egocentric Videos
- 在谷歌Sanpo数据集上精细重标注,引入深度信息融合输入。
- 深度感知使行人与障碍物检测准确率显著提升。
- 适合关注自动驾驶、机器人导航中空间理解的开发者。
长时第一视角视频因视角漂移和缺乏持续几何上下文,给视觉导航带来巨大挑战。尽管近期视觉语言模型在图像和短视频推理中表现优异,其在长第一视角序列中的空间推理能力仍有限。本文研究显式空间信号对基于视觉语言模型的视频理解的影响,不修改模型结构或推理流程。我们引入Sanpo-D,对Google Sanpo数据集进行细粒度重标注,并在面向导航的空间查询任务上基准测试多个VLM。为进一步考察输入层面归纳偏置,我们融合深度图与RGB帧,评估其对空间推理的影响。结果表明,在通用准确率与空间专精之间存在权衡,深度感知和空间锚定表示可显著提升安全关键任务(如行人与障碍物检测)的表现。
原文摘要 · Abstract (English)
Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and short-video reasoning, their spatial reasoning capability in long egocentric sequences remains limited. In this work, we study how explicit spatial signals influence VLM-based video understanding without modifying model architectures or inference procedures. We introduce Sanpo-D, a fine-grained re-annotation of the Google Sanpo dataset, and benchmark multiple VLMs on navigation-oriented spatial queries. To examine input-level inductive bias, we further fuse depth maps with RGB frames and evaluate their impact on spatial reasoning. Our results reveal a trade-off between general-purpose accuracy and spatial specialization, showing that depth-aware and spatially grounded representations can improve performance on safety-critical tasks such as pedestrian and obstruction detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。