arXiv:2512.10071cs.RO2025-12被引 8

提出高效解决方案,在2025行为挑战赛中获亚军,显著超越现有方法。

Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge

  • 基于π₀.₅模型,通过系统性实验优化训练技巧与数据策略。
  • 验证集得分达0.345,大幅领先此前最优水平。
  • 分享实用经验,助力大模型落地复杂具身任务场景。

2025 BEHAVIOR挑战赛旨在通过模拟环境中的物理智能体,严格追踪长时程任务求解进展。BEHAVIOR-1K聚焦日常家庭任务,涵盖人们最希望机器人协助的场景,引入真实环境中复杂的长时程移动操作挑战,弥合当前研究与面向人类的真实应用之间的差距。本报告介绍我们针对2025 BEHAVIOR挑战赛的解决方案,在比赛中取得第二名的优异成绩,并显著优于其他提交方案。基于$π_{0.5}$,我们通过系统性研究训练技术与数据的影响,开展细致的消融实验,揭示预训练与微调阶段均存在显著的缩放收益,最终在验证集上获得0.345的Q分数,远超此前的最先进性能。我们总结了实践经验与设计建议,希望能为整个具身人工智能社区在将强大基础模型适配至复杂具身场景时提供可操作的洞见。项目主页:https://github.com/mli0603/openpi-comet

原文摘要 · Abstract (English)

The 2025 BEHAVIOR Challenge is designed to rigorously track progress toward solving long-horizon tasks by physical agents in simulated environments. BEHAVIOR-1K focuses on everyday household tasks that people most want robots to assist with and these tasks introduce long-horizon mobile manipulation challenges in realistic settings, bridging the gap between current research and real-world, human-centric applications. This report presents our solution to the 2025 BEHAVIOR Challenge in a very close 2nd place and substantially outperforms the rest of the submissions. Building on $π_{0.5}$, we focus on systematically building our solution by studying the effects of training techniques and data. Through careful ablation studies, we reveal the scaling benefits in both the pre-training and post-training phases, leading to a validation Q-score of 0.345, significantly surpassing previous state-of-the-art performance. We summarize our practical lessons and design recommendations that we hope will provide actionable insights for the broader embodied AI community when adapting powerful foundation models to complex embodied scenarios. Project page: https://github.com/mli0603/openpi-comet

具身智能长时程任务机器人挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。