解耦姿态与语言路径,精准定位机器人对话中的物体指代
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution

- 采用分离式融合架构,姿态与文本路径无共享参数
- 在自然手势交互数据集上达31.9% top-1准确率,优于纯姿态或纯文本方案
- 通过可学习门控机制识别语义融合可靠性,适合研究视觉-语言对齐的学者
机器人理解‘把杯子放那一个’这类指令需融合手势、语言与场景几何,但现有3D定位基准存在局限:描述为事后编写,手势为模板化,指向行为受相机视角限制。本研究利用MM-Conv数据集,该数据集来自双人虚拟现实交互,包含完整身体动作捕捉与3D场景图。采用解耦的后期融合架构,使姿态与文本路径不共享任何学习参数,便于通过受控消融实验分离类别、姿态与文本的贡献。使用冻结的MiniLM类别嵌入进行融合,在所有参考类型上均优于仅用姿态或仅用文本的方案,最高达到31.9% top-1准确率。学习到的标量门控会根据文本路径是否具备类别信息,动态切换至相反策略。该机制提供可靠性诊断:若未解耦路径,语义定位的准确性可能源于类别表征偏差而非真正语义对齐。
原文摘要 · Abstract (English)
A robot resolving ``put the cup on that one'' must fuse gesture, language, and scene geometry, yet 3D grounding benchmarks only partially capture this regime: descriptions are written post-hoc, gestures are templated, or pointing is staged for the camera. MM-Conv captures natural co-speech gesture from dyadic VR interaction alongside full-body motion capture and 3D scene graphs. We use it to evaluate pose-language fusion with a decoupled late-fusion architecture in which pose and text pathways share no learned parameters. The two choices together make category, pose, and text contributions easier to isolate through controlled ablations. Fusion with frozen MiniLM category embeddings exceeds pose alone and the best text-only pathway on every reference type, reaching 31.9% top-1. The learned scalar gate flips between opposing policies depending on whether the text pathway has category access. This is a reliability diagnostic: fusion-accuracy claims for semantic grounding systems are indistinguishable from category-representation artifacts unless pathways are architecturally decoupled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。