用框选指令提升扩散模型语义操作泛化能力,发现数据量增长有饱和规律。
Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation
- 用框选目标替代文本指令,结合检测与扩散模型实现精准操控
- 6400条数据验证:性能随数据增长先升后平,存在饱和拐点
- 优先收集多样化物体数据,可使复杂场景成功率达85%
基于扩散的策略在语义操作中泛化能力差,主要因纯文本指令难以在杂乱动态场景中准确引导策略定位目标。本文改用边界框指令直接指定目标,研究性能随数据量的变化规律。为此,构建了Label-UMI手持分割设备及自动化标注流程,高效获取语义标注示范数据;提出一种语义-运动解耦框架,将目标检测与框选引导的扩散策略结合,并引入首帧锚定机制以应对漏检和噪声框。在四个真实任务上验证,使用6400条示范数据,发现泛化性能遵循有界且趋稳的数据缩放规律,表现为边际收益递减。进一步提出以物体多样性为优先的数据采集策略,在复杂场景中实现85%的成功率。所有数据与代码将公开。
原文摘要 · Abstract (English)
Diffusion-based policies generalize poorly in semantic manipulation, a key obstacle to real-world deployment, because text-only instructions cannot reliably steer the policy toward the target object in cluttered, dynamic scenes. We instead use bounding-box instructions to specify the target directly, and study how performance scales with data. To this end, we build Label-UMI, a handheld segmentation device with an automated annotation pipeline for efficiently collecting semantically labeled demonstrations, and propose a semantic-motion-decoupled framework that couples object detection with a bounding-box-guided diffusion policy; a first-frame anchoring mechanism keeps execution robust to missed detections and noisy boxes. We find that generalization follows a bounded, saturating data-scaling law with diminishing returns, validated on four real-world tasks with 6,400 demonstrations, and distill an object-diversity-first collection strategy reaching 85\% success in cluttered scenes. All data and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。