训练可靠机器人奖励模型,需收集失败与危险行为数据。
Position: Good Embodied Reward Models Need Bad Behavior Data

- 用真实失败行为数据训练奖励模型,避免过度奖励危险操作。
- 仅用成功数据会导致模型误判危险行为为合理。
- 适合机器人安全评估与奖励模型优化的研究者参考。
本文主张,构建可靠的具身奖励模型必须投入资源收集“坏”机器人数据:失败、低效、出错甚至危险的行为。尽管奖励模型是基础模型生命周期的核心,当前的具身奖励模型主要基于成功行为训练。我们分析了三种先进具身奖励模型,发现它们系统性地高估了人类评估者会惩罚的行为,包括不安全交互、执行不佳及仅表面满足任务的捷径策略。这些缺陷源于关键数据缺口:负向具身数据成本高、常被过滤或未收录于现有机器人数据集。此外,少量真实坏行为数据即可显著提升与人类偏好的对齐度,并减少昂贵的误报。因此,我们呼吁具身人工智能社区整理并公开坏数据,开发合成坏数据生成工具,建立去中心化物理评估系统,并设计细粒度的具身奖励模型评估基准。
原文摘要 · Abstract (English)
This position paper argues that to obtain reliable embodied reward models, the community must invest in ``bad'' robot data: failed, suboptimal, error-prone, and even hazardous behaviors. While reward models are central to any foundation model's lifecycle, today's embodied reward models are trained primarily on successful behaviors. We analyze three state-of-the-art embodied reward models and find that they systematically over-reward behaviors that real human evaluators would penalize, including unsafe interactions, poor execution, and shortcut strategies that only superficially satisfy tasks. We attribute these failures to a key data gap: the scarcity of negative embodied data which is costly to collect and often filtered out or withheld in existing robotics datasets. Furthermore, we show that even modest exposure to real bad behavior data can improve alignment with human preferences and reduce costly false positives. We therefore call on the embodied AI community to curate and release their bad robot data, build synthetic bad data generation engines, develop more decentralized physical evaluation systems, and design benchmarks for fine-grained embodied reward model evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。