无需训练即可实现3D物体交互功能分割,让机器人更懂人类指令。
3D-TAFS: A Training-free Framework for 3D Affordance Segmentation
- 用大模型融合2D/3D视觉与语言理解,不需训练直接推理。
- 在包含9248张图的室内交互基准上表现优异。
- 适合做智能机器人交互、具身智能系统研发者参考。
将高层语言指令转化为物理世界中的精确机器人动作仍具挑战性,尤其是在考虑与3D物体交互可行性时。本文提出3D-TAFS,一种全新的免训练多模态3D功能分割框架。为支持对这类框架的全面评估,我们构建了IndoorAfford-Bench,一个大规模基准,涵盖6个区域的20种不同室内场景,共9,248张图像,支持标准化交互查询。该框架将大型多模态模型与专用3D视觉网络结合,实现2D与3D视觉理解、语言理解的无缝融合。在IndoorAfford-Bench上的大量实验验证了3D-TAFS在多种场景下处理交互式3D功能分割任务的能力,各项指标表现具有竞争力。结果表明,3D-TAFS有望提升复杂室内环境中基于功能理解的人机交互能力,推动更直观高效的机器人系统在真实场景中的应用。
原文摘要 · Abstract (English)
Translating high-level linguistic instructions into precise robotic actions in the physical world remains challenging, particularly when considering the feasibility of interacting with 3D objects. In this paper, we introduce 3D-TAFS, a novel training-free multimodal framework for 3D affordance segmentation. To facilitate a comprehensive evaluation of such frameworks, we present IndoorAfford-Bench, a large-scale benchmark containing 9,248 images spanning 20 diverse indoor scenes across 6 areas, supporting standardized interaction queries. In particular, our framework integrates a large multimodal model with a specialized 3D vision network, enabling a seamless fusion of 2D and 3D visual understanding with language comprehension. Extensive experiments on IndoorAfford-Bench validate the proposed 3D-TAFS's capability in handling interactive 3D affordance segmentation tasks across diverse settings, showcasing competitive performance across various metrics. Our results highlight 3D-TAFS's potential for enhancing human-robot interaction based on affordance understanding in complex indoor environments, advancing the development of more intuitive and efficient robotic frameworks for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。