arXiv:2511.14161cs.ROcs.CV2025-11被引 3

构建首个支持移动与语言指令的3D家居整理基准,评估机器人综合能力。

RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action

  • 基于3D高斯溅射生成500个逼真家居场景,支持视觉-语言-动作联合训练
  • 提供6400条操作轨迹和1500条导航轨迹,覆盖500个物体与容器
  • 首个端到端真实世界部署的家居整理评测平台,适合具身智能研究者

家居整理是重要应用领域,但现有基准无法建模用户偏好、缺乏移动能力且泛化性差,难以全面评估语言到动作的综合能力。为此,我们提出RoboTidy,一个统一的语言引导家居整理基准,支持视觉-语言-动作(VLA)与视觉-语言-导航(VLN)训练与评估。RoboTidy提供500个逼真3D高斯溅射(3DGS)家居场景(涵盖500个物体与容器),包含碰撞关系,将整理任务定义为“动作(物体,容器)”列表,并提供6400条高质量操作示范轨迹和1500条导航轨迹,支持少样本与大规模训练。我们还在真实世界部署RoboTidy实现物体整理,建立端到端家居整理基准。该平台提供了可扩展的评测环境,填补了具身智能中的关键空白,实现了对语言引导机器人的整体与真实评估。

原文摘要 · Abstract (English)

Household tidying is an important application area, yet current benchmarks neither model user preferences nor support mobility, and they generalize poorly, making it hard to comprehensively assess integrated language-to-action capabilities. To address this, we propose RoboTidy, a unified benchmark for language-guided household tidying that supports Vision-Language-Action (VLA) and Vision-Language-Navigation (VLN) training and evaluation. RoboTidy provides 500 photorealistic 3D Gaussian Splatting (3DGS) household scenes (covering 500 objects and containers) with collisions, formulates tidying as an "Action (Object, Container)" list, and supplies 6.4k high-quality manipulation demonstration trajectories and 1.5k naviagtion trajectories to support both few-shot and large-scale training. We also deploy RoboTidy in the real world for object tidying, establishing an end-to-end benchmark for household tidying. RoboTidy offers a scalable platform and bridges a key gap in embodied AI by enabling holistic and realistic evaluation of language-guided robots.

具身智能3D生成多模态机器人任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。