用人物轨迹统一多视频,实现跨视频问答的高效推理。
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
- 以人物特征为锚点,构建跨视频的层级关系结构。
- 在跨视频问答中达成71.93%人物识别准确率,显著领先现有方法。
- 适合需要跨摄像头分析行为与事件的安防、监控场景。
跨视频问答在传统单视频理解基础上面临更大挑战,尤其体现在跨视频流间建立有效关联及管理多源信息检索的复杂性。本文提出VideoForest框架,通过人物锚定的分层推理机制应对这些挑战。该方法利用人物级特征作为视频间的自然连接点,无需端到端训练即可实现有效的跨视频理解。核心创新包括:1)基于ReID与追踪算法的人物锚定特征提取,建立多视频源间的鲁棒时空关联;2)围绕人物轨迹构建多粒度跨度树结构,分层组织视觉内容;3)多智能体推理框架,高效遍历该层次结构以回答复杂跨视频问题。为评估方法,我们构建了专为人物中心跨视频分析设计的CrossVideoQA基准数据集。实验结果表明,VideoForest在跨视频推理任务中表现优异,人物识别准确率达71.93%,行为分析达83.75%,摘要与推理任务达51.67%,显著超越现有方法。本工作通过人物级特征统一多视频流,建立了一种新型跨视频理解范式,实现分布式视觉信息上的复杂推理,同时保持计算效率。
原文摘要 · Abstract (English)
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning. Our approach leverages person-level features as natural bridge points between videos, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex cross-video queries. To evaluate our approach, we develop CrossVideoQA, a comprehensive benchmark dataset specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning, significantly outperforming existing methods. Our work establishes a new paradigm for cross-video understanding by unifying multiple video streams through person-level features, enabling sophisticated reasoning across distributed visual information while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。