测试AI在多图中通过提问精准找目标的能力。
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
- 设计多轮问答机制,让模型从相似图像中逐步缩小目标范围。
- 在真实任务中实现92%识别成功率,且能有效应对反馈噪声。
- 适合研究具身智能、多轮推理与视觉定位的学者使用。
代理型视觉-语言模型越来越多地通过长时间交互完成任务,但现有评估仍集中于单图单轮的正确性。我们提出AMIGO(Agentic Multi-Image Grounding Oracle Benchmark),一个面向图像画廊中隐藏目标识别的长周期基准。在AMIGO中,系统私密选定一张目标图,模型需通过一系列聚焦属性的“是/否/不确定”问题序列来恢复目标,并遵循严格协议,无效操作将被标记为“跳过”。该设定考验模型在不确定性下的问题选择能力、多轮中约束的一致性追踪能力,以及证据累积后的细粒度判别能力。同时,支持可控的不完美提示,用于探测模型在不一致反馈下的鲁棒性与验证行为。我们在‘猜我最爱的裙子’任务中实例化了AMIGO,报告了包括识别成功率、证据验证、效率、协议合规性、抗噪能力及轨迹级诊断在内的多项指标。
原文摘要 · Abstract (English)
Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identification over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No/Unsure questions under a strict protocol that penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (ii) consistent constraint tracking across turns, and (iii) fine-grained discrimination as evidence accumulates. AMIGO also supports controlled oracle imperfections to probe robustness and verification behavior under inconsistent feedback. We instantiate AMIGO with Guess My Preferred Dress task and report metrics covering both outcomes and interaction quality, including identification success, evidence verification, efficiency, protocol compliance, noise tolerance, and trajectory-level diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。