arXiv:2503.03556cs.CVcs.RO2025-03TPAMI被引 6

轻量模型实现通用物体功能推理,适合本地部署机器人操作。

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

  • 引入新模块增强多模态理解,提升跨场景泛化能力。
  • 在119k图像数据集上达到12.1%性能提升,参数仅187M。
  • 推理速度比GPT-4V快近50倍,适合边缘设备部署。

物体功能推理(affordance reasoning)是任务导向型规划与活动的基础,依赖对物体物理属性和功能的常识性理解,超越简单识别。当前感知驱动的功能推理模型泛化能力不足,限制其在新场景的应用。同时,具备推理能力的大语言模型(LLMs)难以在本地设备部署。为此,我们构建了包含1,496个任务和119,000张图像的大型数据集LVIS-Aff,以提升泛化能力。基于此,提出端到端可训练的Afford-X模型,融合动词注意力与双路融合模块,显著增强多模态理解。该模型相比非LLM方法最佳结果提升12.1%,较此前会议论文提升1.2%;模型仅含187M参数,推理速度接近GPT-4V API的50倍。实验验证其在多种任务与环境中的有效性,证明其在真实世界机器人任务中具有高效与广泛应用潜力。

原文摘要 · Abstract (English)

Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and activities in both humans and Artificial Intelligence (AI). This capability, required for planning and executing daily activities in a task-oriented manner, relies on commonsense knowledge of object physics and functionalities, extending beyond simple object recognition. Current computational models for affordance reasoning from perception lack generalizability, limiting their applicability in novel scenarios. Meanwhile, comprehensive Large Language Models (LLMs) with emerging reasoning capabilities are challenging to deploy on local devices for task-oriented manipulations. Here, we introduce LVIS-Aff, a large-scale dataset comprising 1,496 tasks and 119k images, designed to enhance the generalizability of affordance reasoning from perception. Utilizing this dataset, we develop Afford-X, an end-to-end trainable affordance reasoning model that incorporates Verb Attention and Bi-Fusion modules to improve multi-modal understanding. This model achieves up to a 12.1% performance improvement over the best-reported results from non-LLM methods, while also demonstrating a 1.2% enhancement compared to our previous conference paper. Additionally, it maintains a compact 187M parameter size and infers nearly 50 times faster than the GPT-4V API. Our work demonstrates the potential for efficient, generalizable affordance reasoning models that can be deployed on local devices for task-oriented manipulations. We showcase Afford-X's effectiveness in enabling task-oriented manipulations for robots across various tasks and environments, underscoring its efficiency and broad implications for advancing robotics and AI systems in real-world applications.

功能推理机器人操作轻量模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。