用多段视频生成有视频证据支持的维基式文章
WikiVideo: Article Generation from Multiple Videos
- 通过视频与推理模型协作,从多视频中提炼事件高层语义
- 在真实事件文章生成任务中,新方法显著优于现有模型
- 适合需要多媒体证据支撑的内容生成场景
我们提出一个新任务:基于多个关于现实事件(如自然灾害、政治选举)的多样化视频,生成类似维基百科的文章,且文中所有信息均有视频证据支持。视频是检索增强生成(RAG)的理想输入,但当前RAG主要依赖文本,而视频摘要方法仅关注低层视觉特征,缺乏对高层事件语义的理解。为此,我们构建了WikiVideo基准数据集,包含专家撰写的文章和密集标注的视频证据,支持将视频融入RAG流程。我们进一步提出协同文章生成(CAG)方法,通过一个r1风格的推理模型与VideoLLM的迭代交互,实现比单一VideoLLM更深层次的事件推断。在理想检索和真实RAG设置下,我们评估了主流VideoLLMs与CAG,结果表明CAG持续领先,展现出未来研究的重要方向。
原文摘要 · Abstract (English)
We introduce the task of grounded article generation with the goal of creating a Wikipedia-style article from multiple diverse videos about real-world events -- from natural disasters to political elections -- where all the information in the article is supported by video evidence. Videos are intuitive sources for retrieval-augmented generation (RAG), but most contemporary RAG workflows focus heavily on text while existing methods for video-based summarization focus on low-level scene understanding rather than high-level event semantics. To close this gap, we introduce WikiVideo, a benchmark consisting of expert-written articles and densely annotated videos that provide evidence for articles' claims, facilitating the integration of video into RAG pipelines and enabling the creation of in-depth content that is grounded in multimodal sources. We further propose Collaborative Article Generation (CAG), a novel interactive method for article creation from multiple videos. CAG leverages an iterative interaction between an r1-style reasoning model and a VideoLLM to draw higher-level inferences about the target event than is possible with VideoLLMs alone, which fixate on low-level visual features. We benchmark state-of-the-art VideoLLMs and CAG in both oracle retrieval and RAG settings and find that CAG consistently outperforms alternative methods, while suggesting intriguing avenues for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。