让智能体在多摄像头中推理谁、在哪、何时出现,应对信息不全的挑战。
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

- 用智能体交互方式解决多摄像头下人员搜索难题,需规划提问与工具使用。
- 在真实场景中测试,最先进方法在时间推理上准确率仅59.0%。
- 引入时空拓扑图辅助决策,适合安全监控与司法调查场景研究者。
我们提出ARGOS,首个将多摄像头人员搜索重构为需智能体进行规划、提问与排除候选者的交互式推理任务的基准与框架。该智能体接收模糊目击陈述,须决定提问内容、调用空间或时间工具的时机,并解释模糊回应,且受限于有限对话轮次。推理基于编码摄像头连接关系与实证验证过渡时间的时空拓扑图(STTG)。基准包含14个真实场景中的2,691个任务,分为三个递进赛道:语义感知(谁)、空间推理(哪)和时间推理(何时)。实验使用四种LLM主干模型显示,该任务仍未解决(赛道2最佳TWS为0.383,赛道3为0.590),消融实验表明移除领域专用工具会使准确率下降最多达49.6个百分点。
原文摘要 · Abstract (English)
We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。