给机器人操作反馈加了个智能质检员,让新手更快练出好数据。
Closing the Loop in Teleoperation: Episode-Level Data Quality Assessment and Feedback for High-Quality Demonstration Collection

- 用任务进展和机械数据自动评估操作质量
- 新手用反馈后更快产出高质量示范数据
- 适合需要大量真实操作数据的机器人训练场景
工业自动化正迎来关键转折点,物理人工智能推动系统从固定程序向更灵活自适应方向演进。这一转变催生了对大规模真实机器人示范数据的巨大需求,使远程操控成为数据采集的重要手段。然而实践中高质量的远程操控示范仍难获得,新手操作者常完成任务但动作低效、反复修正或接近关节极限,影响后续使用效果。本文提出数据质量评估与反馈(DQAF)框架,通过语义任务进度和机器人遥测数据实现远程操控闭环反馈。该框架提取子任务进度、运动平滑性、停滞、运动学限制等质量信号,转化为结构化评估和可操作的自然语言建议。不同于简单的成功/失败判断,本系统解释为何不达标并指出需改进的具体行为。通过诊断验证与试点用户研究验证:在数据筛选中,系统生成的拒因与改进建议与人工评审一致;三名新手在两个操作任务中,接收即时反馈的用户进步更快,更早产出高质量示范数据。
原文摘要 · Abstract (English)
Industrial automation is at a pivotal moment, as Physical AI is driving a transition from rigid, hand-engineered automation systems toward more flexible and adaptive systems. This shift has created a growing demand for large-scale, real-world robot demonstration data, making teleoperation an increasingly important mechanism for data collection. However, high-quality teleoperated demonstrations remain difficult to obtain in practice, as novice operators often produce episodes that are task-successful but suboptimal for downstream use due to inefficient motion, repeated corrections, or operation near robot joint limits. We present a Data Quality Assessment and Feedback (DQAF) framework that closes the loop in teleoperation by providing immediate post-episode feedback grounded in semantic task progress and robot telemetry. The framework extracts quality relevant signals such as sub-task progress, motion smoothness, stalls, kinematic limits and converts them into structured quality assessments and actionable natural-language feedback. Unlike binary success or failure feedback, the proposed system explains why an episode is suboptimal and highlights specific behaviors to correct in the next trial. We evaluate the framework through a diagnostic validation study and a pilot user study. In the validation study, the system is compared with a human reviewer during dataset curation, producing rejection reasons and actionable feedback for improvement. In the pilot study with three novice operators across two manipulation tasks, the operator who received the systems immediate, automated post-episode feedback improved faster than those who did not, producing higher-quality demonstrations sooner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。