arXiv:2608.17453cs.RO2026-08

让机器人用双眼看世界,智能选择有用视觉信息。

EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

论文配图:EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control
图 1 · 摘自论文原文
  • 通过路由机制保留主视角,生成对齐的辅助视角特征
  • 实测完成率60%,抓取成功率100%,严重遮挡下恢复率达80%
  • 适合需要精准空间感知的复杂人形机器人任务

长时序人形机器人视觉-语言-动作控制需利用头戴双目相机的互补视图,同时兼容预训练模型。现有方法常丢弃互补立体信息或融合额外观测但未保留原始主视角路径,也未根据机器人本体状态适配辅助信息。我们提出EATR-Stereo,一种具身感知的令牌路由框架:保留主视角令牌,并通过查询同步的辅助视角令牌序列构建与主视角对齐的跨视图辅助令牌(CVAT)。一个身体分段的本体感觉编码器进一步基于机器人配置历史,实现对辅助信息使用的细粒度调控。路由后的辅助流在不改变预训练视觉-语言模型的前提下,增强语言与主视觉上下文。在33自由度物理人形机器人(37维本体感觉状态)上,我们在超过100秒的搜索-接近-抓取-放置-返回任务中测试了九种配置。EATR-Stereo实现60.0%全任务成功率、100.0%抓取成功率和80.0%阶段成功率。在严重不对称遮挡下,恢复率提升至80%,远超仅使用CVAT的30%。消融实验表明,保留主视角令牌以及将跨视图辅助特征与结构化本体感觉路由结合至关重要。结果证明,有选择地路由成对立体证据可显著提升长期人形机器人控制的空间定位可靠性。

原文摘要 · Abstract (English)

Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary information to robot embodiment. We present EATR-Stereo, an embodiment-aware token-routing framework that retains primary-view tokens and constructs primary-aligned Cross-View Auxiliary Tokens (CVATs) by querying the synchronized auxiliary-view token sequence. A body-segmented proprioceptive encoder further conditions token-wise auxiliary usage on robot configuration history, enabling selective incorporation of stereo evidence during action generation. The routed auxiliary stream augments the language and primary-visual context of a pretrained VLA while keeping its vision--language model frozen. On a 33-DoF physical humanoid with a 37-D proprioceptive state, we evaluate nine configurations in over-100-s search--approach--grasp--place--return tasks. EATR-Stereo achieves 60.0% full-task success, 100.0% grasp success, and 80.0% stage success. Under severe asymmetric occlusion, it improves recovery to 80% compared with 30% for CVAT alone. Ablation studies further show the importance of preserving primary tokens and combining cross-view auxiliary features with structured proprioceptive routing. These results demonstrate that selectively routed paired stereo evidence improves spatial grounding for reliable long-horizon humanoid VLA control.

人形机器人立体视觉多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。