arXiv:2608.07417cs.CVcs.AI2026-08中稿 · ACM Multimedia 202…

让AI通过照片定位视频中的人,实现跨镜头跟踪与行为理解。

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

论文配图:I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
图 1 · 摘自论文原文
  • 用参考图像引导模型理解视频中特定人物的行动和关系。
  • 构建1377个复杂视频数据集,涵盖从识人到因果推理的六级难度。
  • 适合研究视频理解、人物追踪及多模态大模型的开发者使用。

现实世界中的视频推理常涉及多模态、多源输入,而现有任务多限于简化视频-文本场景,难以实现身份匹配与以人为中心的推理。为此,我们提出身份条件查询(ICQ)任务,要求模型联合关联输入视频与某个人物的参考图像,并利用该条件解决身份定位、行为理解、时间推理等挑战。基于此,我们构建ISYV(I Seek You in Videos)系统,包含三部分:(1) ISYV-Bench,一个包含1,377个真实复杂视频和1,377个问答对的评估基准,分为六级难度,覆盖从身份识别到因果推理的能力;(2) ISYV-75K,通过自动化标注、多阶段验证与人工审核构建的75,000条高质量训练样本;(3) ISYV-Framework,含面向ICQ的模型与训练策略,可在无需额外帧级标注的情况下学习利用关键视频片段。大量实验表明,主流闭源与开源多模态大模型在ISYV-Bench上表现不佳,尤其在跨域身份匹配与长时程追踪方面。ISYV-Model优于多数基线,在某些方面接近闭源模型性能。总体而言,ISYV提供统一的任务定义、可扩展的数据集与基准,以及建模洞见,推动以人为中心的视频推理发展。

原文摘要 · Abstract (English)

Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges. Building on ICQ, we present ISYV (I Seek You in Videos), a systematic solution comprising three components: (1) ISYV-Bench, a challenging evaluation benchmark with 1,377 real-world complex videos and 1,377 question-answer pairs, organized into six difficulty levels spanning capabilities from identity recognition to causal reasoning; (2) ISYV-75K, a large-scale training set of 75K high-quality samples constructed via automated annotation, multi-stage verification, and manual review; and (3) ISYV-Framework, containing an ICQ-oriented model and training strategy for learning to exploit informative video shots without additional shot-level annotations. Extensive experiments show that both mainstream closed-source and open-source MLLMs struggle on ISYV-Bench, especially in cross-domain identity matching and long-horizon tracking. ISYV-Model outperforms strong baselines and in some aspects approaches closed-source performance. Overall, ISYV provides a unified task definition, scalable datasets/benchmarks, and modeling insights for person-centric video reasoning.

视频理解人物追踪多模态身份识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。