arXiv:2604.12762cs.CVcs.AI2026-04中稿 · CVPR

让智能体在多摄像头中推理谁、在哪、何时出现,应对信息不全的挑战。

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

论文配图:ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
图 1 · 摘自论文原文
  • 用智能体交互方式解决多摄像头下人员搜索难题,需规划提问与工具使用。
  • 在真实场景中测试,最先进方法在时间推理上准确率仅59.0%。
  • 引入时空拓扑图辅助决策,适合安全监控与司法调查场景研究者。

我们提出ARGOS,首个将多摄像头人员搜索重构为需智能体进行规划、提问与排除候选者的交互式推理任务的基准与框架。该智能体接收模糊目击陈述,须决定提问内容、调用空间或时间工具的时机,并解释模糊回应,且受限于有限对话轮次。推理基于编码摄像头连接关系与实证验证过渡时间的时空拓扑图(STTG)。基准包含14个真实场景中的2,691个任务,分为三个递进赛道:语义感知(谁)、空间推理(哪)和时间推理(何时)。实验使用四种LLM主干模型显示,该任务仍未解决(赛道2最佳TWS为0.383,赛道3为0.590),消融实验表明移除领域专用工具会使准确率下降最多达49.6个百分点。

原文摘要 · Abstract (English)

We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.

多摄像头智能体人物搜索时空推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。