提出跨视角点对应任务,提升视觉语言模型的精细空间交互能力。
Towards Cross-View Point Correspondence in Vision-Language Models
- 构建分层评估基准,模拟人类感知-推理-对应认知过程。
- 在378K数据集上训练模型,相比Gemini-2.5-Pro提升39.7%准确率。
- 聚焦可操作区域,适合研究具身智能与精准交互的开发者。
跨视角对应是空间理解与具身智能的基础能力,但当前视觉语言模型在实现精确点级对应方面仍严重不足,而这对于精准功能交互至关重要。为此,我们提出了跨视角点对应(CVPC)任务,并构建了跨视图点对应基准(CrossPoint-Bench),其设计灵感源于人类“感知-推理-对应”的认知过程。评估显示,当前最先进的模型(如Gemini-2.5-Pro)与人类表现仍有超过54.65%的准确率差距,暴露出从粗粒度判断向细粒度坐标预测过渡的挑战。为解决该问题,我们构建了涵盖900个场景、包含378,000个问答对的CrossPoint-378K数据集,聚焦真实世界中的可操作区域,以更贴近实际交互场景。此外,我们提出CroPond模型,其在CrossPoint-Bench上达到最优性能,相比Gemini-2.5-Pro提升39.7%准确率,为未来跨视角对应研究奠定基础。相关基准、数据集与模型已开源:https://github.com/WangYipu2002/CrossPoint。
原文摘要 · Abstract (English)
Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), especially in achieving precise point-level correspondence, which is crucial for precise affordance interaction. So we propose the Cross-View Point Correspondence (CVPC) task and CrossPoint-Bench, a comprehensive benchmark with hierarchical design, inspired by the human cognitive process of "perceive", "reason", and "correspond". Our evaluation shows the state-of-the-art models (e.g., Gemini-2.5-Pro) still fall far behind humans, with a gap of over 54.65% in overall accuracy, exposing a challenge in transitioning from coarse-grained judgement to fine-grained coordinate prediction. To address this problem, we construct CrossPoint-378K, a dataset with 378K question-answering pairs across 900 scenes, focused on actionable affordance regions that better reflect real-world manipulation and interaction scenarios. Furthermore, we propose CroPond that trained on the CrossPoint-378K dataset. Our CroPond achieves state-of-the-art performance on CrossPoint-Bench, surpassing Gemini-2.5-Pro by 39.7% accuracy, which offers a foundation for advancing future work on cross-view correspondence. The benchmark, dataset, and model are publicly available at https://github.com/WangYipu2002/CrossPoint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。