基于视觉-语言-动作模型,实现家庭场景下复杂任务的高效精准执行。
Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- 引入相关噪声流匹配,提升训练效率并生成连贯动作序列。
- 在50个任务中达成26%的q-score,跨公开与私有榜单表现优异。
- 适合需要多臂操作与环境感知的机器人任务研究者参考。
我们提出一种视觉-动作策略,在2025 BEHAVIOR挑战赛中获得第一名。该挑战赛是大规模基准测试,包含50个多样化的长时序家庭任务,运行于照片级真实感仿真环境中,要求双臂操作、导航及上下文感知决策。基于Pi0.5架构,我们引入多项创新:采用相关噪声进行流匹配,提升训练效率,并支持关联感知的补全以生成平滑动作序列;应用可学习的混合层注意力与系统2阶段追踪机制,用于歧义消解。训练阶段采用多样本流匹配降低方差,推理阶段则结合动作压缩与特定挑战修正规则。该方法在公开与私有排行榜上均实现26%的q-score,覆盖全部50个任务。
原文摘要 · Abstract (English)
We present a vision-action policy that won 1st place in the 2025 BEHAVIOR Challenge - a large-scale benchmark featuring 50 diverse long-horizon household tasks in photo-realistic simulation, requiring bimanual manipulation, navigation, and context-aware decision making. Building on the Pi0.5 architecture, we introduce several innovations. Our primary contribution is correlated noise for flow matching, which improves training efficiency and enables correlation-aware inpainting for smooth action sequences. We also apply learnable mixed-layer attention and System 2 stage tracking for ambiguity resolution. Training employs multi-sample flow matching to reduce variance, while inference uses action compression and challenge-specific correction rules. Our approach achieves 26% q-score across all 50 tasks on both public and private leaderboards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。