arXiv:2607.03693cs.RO2026-07被引 2

让机器人在缺传感器时仍能稳定执行复杂任务

CoRE-VLA: Towards Scalable and Robust Vision-Language-Action Modeling via Conditional Routing of Experts

论文配图:CoRE-VLA: Towards Scalable and Robust Vision-Language-Action Modeling via Conditional Routing of Experts
图 1 · 摘自论文原文
  • 按传感器可用性动态选择专家模块,实现稀疏计算
  • 在多个长序列任务上表现优于现有模型,零样本泛化能力强
  • 适合部署在传感器配置多样的真实机器人场景

视觉-语言-动作(VLA)模型推动了通用机器人操作的发展,但实际部署中面临挑战:机器人传感器配置多样且可能意外失效,不同机体设计也常缺少某些传感器。统一策略需在传感器可用时利用其输入,缺失时仍保持可靠。现有VLA策略通过共享密集计算耦合动作生成与固定传感器集,导致传感器缺失时性能脆弱,难以跨任务和长程行为专精。本文提出CoRE-VLA框架,将动作生成建模为上下文感知的稀疏计算:传感器可用性门控模态专用专家,支持缺失传感器下的渐进退化且无需重训练;任务意图进一步引导动作表示至相关专家,提升任务专精性。虽可适配多种辅助传感器,实验聚焦深度信息这一关键模态。在LIBERO、RoboCasa GR1 Tabletop及双臂实机任务上,CoRE-VLA在长程与多任务基准上均表现优异,超越密集动作生成基线与强预训练VLA模型,包括在未见场景中的零样本泛化能力。模态分析表明,该模型可在有深度信息时有效利用,缺失时仍保持鲁棒。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have advanced generalist robotic manipulation, yet real-world deployment reveals a fundamental challenge: robots are equipped with diverse and heterogeneous sensor configurations, auxiliary sensors can fail unexpectedly during operation, and different robot embodiments often lack certain sensors by design. A unified policy that can exploit auxiliary perceptual inputs when available while remaining reliable under sensor absence, whether incidental or by design, is therefore essential for practical deployment. However, existing VLA policies couple action generation to a fixed sensor set through shared dense computation, making them brittle when sensors are missing and limiting their ability to specialize across diverse tasks and long-horizon behaviors. We propose CoRE-VLA, a scalable and robust VLA framework that formulates action generation as context-conditioned sparse computation. Sensor availability gates modality-specialized experts, enabling graceful degradation under missing sensors without retraining. Task intent further routes action-side representations to task-relevant experts, improving specialization across diverse tasks and long-horizon subgoals. While the framework is designed to accommodate different auxiliary sensors, we focus on depth as a representative and practically important auxiliary modality in our experiments. Experiments on LIBERO, RoboCasa GR1 Tabletop, and real-world dual-arm manipulation show that CoRE-VLA achieves strong results on long-horizon and multi-task benchmarks, and outperforms both a dense-action-generator ablation and a strong pretrained VLA baseline, including in zero-shot generalization to unseen scenarios. Modality analysis shows that CoRE-VLA can exploit auxiliary depth when available while remaining robust when depth is unavailable during deployment.

机器人多模态稀疏计算鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。