提升机器人视觉指代精度,通过智能数据生成与纠错机制实现77.2%准确率
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

- 用代理驱动生成大量带语义标注的候选样本,构建5.5万条处理数据
- 采用可验证的数据流水线,生成1万条真实标注样本并保留备用集
- 引入注意力修正与坐标纠偏模块,适合需要精准定位的具身智能应用
视觉指代将语言指令映射到像素坐标,是具身AI的核心能力。本文提出PointArena 2026方案,在基准测试中达到77.2%整体准确率,排名第二。针对三大失效模式:第一,代理驱动的数据合成构建大规模语义与锚点相对的候选池,服务器库存含55,372条处理输出、53,772个去重样本ID、37,574条可训练完成或接受的行;第二,设计确定性可调节数据流水线,通过掩码、模板和路径验证生成10,000样本主集及备用样本;第三,引入两个模型级模块:AttnRes加入门控跨块注意力提升可引导性,ABC校正利用视觉特征编码扰动坐标实现通用坐标定位。类别感知路由融合互补专家,局部验证显示各类任务准确率分别为:93.9%(可用性)、82.6%(空间关系)、78.2%(推理)、70.4%(计数)、63.0%(可引导性)。
原文摘要 · Abstract (English)
Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and ranks second on the benchmark. The ap proach targets three failure modes. First, agent-driven syn thesis builds large semantic and anchor-relative candidate pools; the server inventory contains 55,372 processed out puts, 53,772 de-duplicated sample IDs, and 37,574 train able completed or accepted rows. Second, a determinis tic steerable-data pipeline creates a verified 10,000-sample main set, plus reserve samples, using masks, templates, and path verification. Third, two model-side modules address complementary errors: AttnRes adds gated cross-block at tention for steerability, while ABC correction encodes per turbed coordinates with visual features for general coordi nate grounding. Category-aware routing combines comple mentary specialists; local validation used to select experts records 93.9% Affordance, 82.6% Spatial Relation, 78.2% Reasoning, 70.4% Counting, and 63.0% Steerability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。