解决点云视觉语言模型的几何幻觉问题,提升3D结构预测准确性。
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment

- 通过几何奖励分配机制,精准定位并优化关键几何令牌。
- 3D KPA提升至0.93,框交并比达0.686,投影一致性达0.852。
- 适合需要高精度3D空间推理的机器人导航与具身智能研究者。
点云视觉语言模型有望赋予具身智能体可执行的空间推理能力,但常出现几何幻觉,即预测的3D结构与观测到的2D现实矛盾。我们发现根本原因并非表征瓶颈,而是强化学习中稀疏几何标记被噪声和全局广播奖励所淹没。为此提出几何奖励信用分配框架,将整体监督解耦为领域特定信号,并仅传递给对应标记段。该机制将模糊反馈转化为精确梯度更新,实现从通用策略优化到结构化对齐的转变。同时引入重投影一致性项,作为跨模态验证器惩罚物理上不可能的几何结构。在基于ShapeNetCore的校准基准上验证,本方法将3D KPA从0.64提升至0.93,3D边界框交并比达0.686,重投影一致性分数升至0.852。关键的是,2D定位性能保持稳定,标志着从合理文本输出迈向可物理验证的空间预测的重要进展。
原文摘要 · Abstract (English)
Point-Vision-Language Models promise to empower embodied agents with executable spatial reasoning, yet they frequently succumb to geometric hallucination where predicted 3D structures contradict the observed 2D reality. We identify a key cause of this failure not as a representation bottleneck but as a structural misalignment in reinforcement learning, where sparse geometric tokens are drowned out by noisy and broadcasted sequence-level rewards. To resolve this causal dilution, we propose Geometric Reward Credit Assignment, a framework that disentangles holistic supervision into field-specific signals and routes them exclusively to their responsible token spans. This mechanism transforms vague feedback into precise gradient updates and effectively turns generic policy optimization into targeted structural alignment. Furthermore, we internalize physical constraints via a Reprojection-Consistency term which serves as a cross-modal verifier to penalize physically impossible geometries. Validated on a calibrated benchmark derived from ShapeNetCore, our approach bridges the reliability gap by boosting 3D KPA from 0.64 to 0.93, increasing 3D bounding box intersection over union to 0.686, and raising reprojection consistency scores to 0.852. Crucially, these gains are achieved while maintaining robust 2D localization performance, marking a meaningful step from plausible textual outputs toward physically verifiable spatial predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。