用3D场景建模+思维链推理,让机器人更懂空间、看得更远。
PRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning
- 融合3D点云与语言模型,减少遮挡和视角偏差。
- 引入思维链推理,提升长期决策能力,偏好判断准确率更高。
- 适合需要精准空间理解的机器人任务,如抓取与导航。
我们提出PRISM,一种新型框架,旨在克服基于2D的偏好强化学习(PBRL)局限性,通过统一3D点云建模与前瞻式偏好优化。核心在于采用3D点云-语言模型(3D-PC-LLM),缓解遮挡与视角偏差,确保偏好信号更稳定且空间一致。同时,利用思维链(CoT)推理引入长时程考量,避免静态偏好对比中的短视反馈。相比传统方法,该框架在偏好一致率、策略收敛速度及未见环境泛化能力上均有显著提升。实证结果涵盖机器人操作与自主导航任务,验证了其在需精确空间理解与可靠长期决策场景中的实际应用潜力。通过融合3D几何感知与基于思维链的偏好建模,PRISM为可扩展的人类对齐强化学习奠定基础。
原文摘要 · Abstract (English)
We propose PRISM, a novel framework designed to overcome the limitations of 2D-based Preference-Based Reinforcement Learning (PBRL) by unifying 3D point cloud modeling and future-aware preference refinement. At its core, PRISM adopts a 3D Point Cloud-Language Model (3D-PC-LLM) to mitigate occlusion and viewpoint biases, ensuring more stable and spatially consistent preference signals. Additionally, PRISM leverages Chain-of-Thought (CoT) reasoning to incorporate long-horizon considerations, thereby preventing the short-sighted feedback often seen in static preference comparisons. In contrast to conventional PBRL techniques, this integration of 3D perception and future-oriented reasoning leads to significant gains in preference agreement rates, faster policy convergence, and robust generalization across unseen robotic environments. Our empirical results, spanning tasks such as robotic manipulation and autonomous navigation, highlight PRISM's potential for real-world applications where precise spatial understanding and reliable long-term decision-making are critical. By bridging 3D geometric awareness with CoT-driven preference modeling, PRISM establishes a comprehensive foundation for scalable, human-aligned reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。