提出新基准与框架,提升机器人操作复杂物品的泛化能力
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
- 用分层基准测试跨部件、跨实例的长程操作挑战
- 在五个环境上实现91.2%成功率,显著超越现有方法
- 适合研究机器人操作与具身智能的开发者
交互式可动物体操作需要长时间、多步骤与电器设备互动并保持物理一致性。现有视觉语言与扩散模型策略在不同部件、实例和类别间泛化能力差。我们首先提出 ArtiBench,一个包含厨房、储物、办公和工具环境的五级基准,支持从跨部件、跨实例变化到长程多物体任务的结构化评估,揭示了可动物体操作的核心泛化难题。基于此基准,我们提出 ArtiBrain,一种模块化框架,统一高层推理与自适应底层控制。ArtiBrain 使用基于 VLM 的任务推理器(GPT-4.1)分解并验证子目标,采用混合控制器结合几何感知关键帧执行与可觉察引导的扩散模型,实现精确且可解释的操作。可觉察记忆库持续积累成功执行经验,并将部件级可操作性传播至未见过的可动部件与配置。在 ArtiBench 上的大量实验表明,ArtiBrain 在鲁棒性和泛化能力上显著优于现有先进多模态与扩散模型方法。代码与数据集将在论文接收后公开。
原文摘要 · Abstract (English)
Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-based policies struggle to generalize across parts, instances, and categories. We first introduce ArtiBench, a five-level benchmark covering kitchen, storage, office, and tool environments. ArtiBench enables structured evaluation from cross-part and cross-instance variation to long-horizon multi-object tasks, revealing the core generalization challenges of articulated object manipulation. Building on this benchmark, we propose ArtiBrain, a modular framework that unifies high-level reasoning with adaptive low-level control. ArtiBrain uses a VLM-based Task Reasoner (GPT-4.1) to decompose and validate subgoals, and employs a Hybrid Controller that combines geometry-aware keyframe execution with affordance-guided diffusion for precise and interpretable manipulation. An Affordance Memory Bank continually accumulates successful execution episodes and propagates part-level actionable affordances to unseen articulated parts and configurations. Extensive experiments on ArtiBench show that our ArtiBrain significantly outperforms state-of-the-art multimodal and diffusion-based methods in robustness and generalization. Code and dataset will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。