为可穿戴设备设计3D空间推理新基准和工具框架,提升定位与测量能力。
R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables
- 构建基于真实视角视频的3D空间推理工具调用框架
- 在15类问题上达73.5%准确率,显著超越现有方法
- 适合做智能助手、机器人视觉理解的研究者参考
从第一人称RGB-D视频中进行定量3D空间推理是下一代可穿戴助手的关键能力。然而,现有基准无法反映自然第一人称视频、带深度信息的输入以及具有挑战性的量化3D空间问答任务。为此,我们提出R3D-Bench(3D推理),一个包含3,033个定量空间推理问题的基准,覆盖15种类型,基于Aria Digital Twin的57段第一人称视频序列构建。为建立强基线,我们引入R3D——一种模型无关的空间工具调用框架。不同于将3D信息直接嵌入模型输入,R3D通过分割和深度提升的对象表示重建3D场景,并通过8个可组合的空间工具向LLM提供信息。在R3D-Bench上,R3D结合Qwen3-VL 235B达到73.5%的平均相对准确率,显著优于最佳深度增强基线(CuTR+Tools,61.9%)和最佳仅图像基线(Gemini 3 Flash,46.5%)。
原文摘要 · Abstract (English)
Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants. Yet existing benchmarks do not reflect the challenges of handling (1) natural egocentric video, (2) posed RGB-D video inputs, and (3) challenging quantitative 3D spatial reasoning Q&A. To fill this gap, we introduce R3D-Bench (Reasoning in 3D), a benchmark of 3,033 quantitative spatial reasoning questions across 15 types -- spanning multiple-choice, distance-based, and volumetric reasoning questions -- built on top of 57 egocentric video sequences from Aria Digital Twin. To set a strong baseline on this dataset, we introduce R3D, a model-agnostic spatial tool-calling framework. In contrast to existing approaches that directly embed 3D information into the model's input representation, R3D constructs a 3D scene from video using segmentation and depth-lifted object representations. It provides this information to an LLM through eight composable spatial tools. On R3D-Bench, R3D with Qwen3-VL 235B achieves 73.5% mean relative accuracy, substantially outperforming the best depth-enabled baseline (CuTR+Tools, 61.9%) and the best RGB-only baseline (Gemini 3 Flash, 46.5%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。