在虚拟现实场景中,通过多视角分析实现对无交互对象状态变化的精准推理。
ObjChangeVR: Object State Change Reasoning from Continuous Egocentric Views in VR Environments
- 融合视角感知与时间检索,定位关键帧并整合多视角信息。
- 在自建数据集上,模型性能超越基线方法超过15个百分点。
- 适合研究视觉语言模型在复杂动态场景中推理能力的学者。
多模态大语言模型(MLLM)为虚拟现实中的自然语言场景变化查询提供了新路径。以往研究聚焦于用户直接交互的自我视角视频,但对象状态变化可能发生在背景中,缺乏明显运动线索,难以检测。此外,该场景尚无基准评测数据集。为此,我们提出ObjChangeVR-Dataset,专门用于评估对象状态变化问答任务。同时设计了ObjChangeVR框架,结合视角感知与时间检索定位相关帧,并通过跨视角推理融合不一致证据。大量实验表明,该框架在多个MLLM上显著优于基线方法。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) offer a promising approach for natural language-based scene change queries in virtual reality (VR). Prior work on applying MLLMs for object state understanding has focused on egocentric videos that capture the camera wearer's interactions with objects. However, object state changes may occur in the background without direct user interaction, lacking explicit motion cues and making them difficult to detect. Moreover, no benchmark exists for evaluating this challenging scenario. To address these challenges, we introduce ObjChangeVR-Dataset, specifically for benchmarking the question-answering task of object state change. We also propose ObjChangeVR, a framework that combines viewpoint-aware and temporal-based retrieval to identify relevant frames, along with cross-view reasoning that reconciles inconsistent evidence from multiple viewpoints. Extensive experiments demonstrate that ObjChangeVR significantly outperforms baseline approaches across multiple MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。