用真实扫描数据提升场景理解,让大模型和机器人更好适应现实环境。
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
- 统一用USD格式整合多样扫描数据,支持不同应用定制
- 大模型场景编辑成功率达80%,机器人仿真策略学习成功率87%
- 适合做具身智能、机器人训练和真实世界场景建模的研究者
真实世界的3D场景扫描数据具有高度真实性,有助于提升下游任务在现实中的泛化能力。然而,数据量庞大、标注格式多样及工具兼容性差等问题限制了其应用。本文提出一种基于USD的统一标注集成方法,并设计面向具体应用的USD变体。我们识别出使用完整真实扫描数据集的关键挑战,并提出相应的缓解策略。所提方法在两个下游任务中验证有效:基于大语言模型的场景编辑任务中,实现了80%的成功率;在机器人仿真中,策略学习的成功率达到87%。
原文摘要 · Abstract (English)
Real-world 3D scene-level scans offer realism and can enable better real-world generalizability for downstream applications. However, challenges such as data volume, diverse annotation formats, and tool compatibility limit their use. This paper demonstrates a methodology to effectively leverage these scans and their annotations. We propose a unified annotation integration using USD, with application-specific USD flavors. We identify challenges in utilizing holistic real-world scan datasets and present mitigation strategies. The efficacy of our approach is demonstrated through two downstream applications: LLM-based scene editing, enabling effective LLM understanding and adaptation of the data (80% success), and robotic simulation, achieving an 87% success rate in policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。