用可验证标注训练模型精准指点,提升视觉语言定位准确率。
PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

- 将框、掩码等标注转为指点指令,用隐藏验证器评分。
- 在PointArena上将准确率从56.11%提升至65.58%。
- 适合需要精确空间定位的机器人和交互系统应用。
视觉语言模型日益依赖点坐标作为紧凑且可执行的界面,用于图形用户界面交互、机器人操作和交互式视觉系统中的视觉定位。然而,由于监督空间本质上不唯一——同一目标区域可能有多个有效坐标——且多实例指令需满足目标覆盖、数量一致性和去重要求,学习可靠指指点行为仍具挑战。本文提出PointRL,一种从异构标注证据中学习点级定位的可验证强化学习框架。PointRL将边界框、掩码和实例标签转换为指点指令,同时保留其目标支持、实例归属和集合约束作为隐藏验证证据(即不放入提示中,由确定性检查器评分)。提出的奖励函数评估可解析性、点有效性、实例覆盖率、基数一致性以及冗余或缺失预测。在PointArena上,PointRL将Qwen3.5-4B的整体准确率从56.11%提升至65.58%。在RoboSpatial、BLINK和Ref-Adv上的进一步评估显示,相同主干模型在外部基准上也取得提升,表明可验证点级反馈可能有助于这些场景下的空间定位。
原文摘要 · Abstract (English)
Vision-language models (VLMs) increasingly rely on point coordinates as a compact and executable interface for visual grounding in GUI interaction, robotic manipulation, and interactive visual systems. However, learning reliable pointing behavior remains difficult because the supervision space is inherently non-unique: many coordinates may be valid within the same target region, while multi-instance instructions require target coverage, count consistency, and duplicate suppression. This work presents PointRL, a verifiable reinforcement learning framework that learns point-level grounding from existing heterogeneous annotation evidence. PointRL converts bounding boxes, masks, and instance labels into pointing instructions, while retaining their target supports, instance membership, and set constraints as hidden verifier evidence, i.e., annotations kept outside the prompt and used by a deterministic checker to score predictions. The proposed reward evaluates parseability, point validity, instance coverage, cardinality consistency, and redundant or missing predictions. On PointArena, PointRL improves the overall accuracy of Qwen3.5-4B from 56.11% to 65.58%. Further evaluations on RoboSpatial, BLINK, and Ref-Adv show same-backbone gains on the evaluated external benchmarks, suggesting that verifiable point-level feedback may benefit spatial grounding in these settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。