提出新训练方法,让机器人在多传感器下自动识别关键信息,提升真实场景适应力。
Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies

- 根据每帧每传感的证据强度,动态调节训练目标,区分有用与无用信号。
- 在传感器被干扰或只剩单一传感器时,成功率最高提升120%。
- 适用于双臂机械手和触觉融合等不同机器人系统,通用性强。
视觉-语言-动作(VLA)策略融合多模态感知输入,但受限于少量同质化机器人示范数据,容易产生虚假的跨传感器相关性,而非任务相关信号,这种现象称为模态纠缠。在真实世界遮挡与干扰下,表现为对无用传感器噪声敏感,以及仅剩一个有效传感器时的单模态不足。本文提出证据门控正则化(EGR),一种无需推理开销的模态无关训练目标。EGR通过每帧每传感器的任务相关性信号,调控两个状态条件一致性目标:低证据传感器保持不变性,高证据传感器确保单模态充分性。我们基于BEHAVIOR-1K构建基准,包含快速推理诊断套件和47个针对模态纠缠的滚动技能评估。在该基准及两个不同构型的真实机器人平台上验证:双臂康奈尔机械臂配三台RGB相机,以及单臂MELFA ASSISTA结合视觉与GelSight触觉传感器。EGR将模拟成功率从12.5%提升至16.4%(+31%),在无用传感器干扰下从9.4%升至16.5%(+75%),单传感器故障时从2.8%增至6.1%(+120%)。在物理对象干扰下,双臂平台成功率从30%升至85%(+183%),触觉平台从55%升至70%(+27%)。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。