arXiv:2504.06272cs.IRcs.AI2025-04被引 1

用AI自动从海量视频中发现跨模态实体,支持个性化搜索与内容发现。

RAVEN: An Agentic Framework for Multimodal Entity Discovery from Large-Scale Video Collections

  • 通过视觉、音频、文本多模态融合,自主理解视频主题和通用实体
  • 动态生成领域专属实体与属性,提升复杂场景的识别能力
  • 支持多种模型接入,适合大规模视频数据的智能检索应用

我们提出RAVEN,一个面向大规模视频集合的多模态实体发现与检索的自适应智能体框架。通过融合视觉、音频和文本信息,RAVEN自主处理视频数据,生成可用于下游任务的结构化、可操作表示。主要贡献包括:(1) 一类主题理解步骤,用于推断视频主题和通用实体;(2) 动态生成领域特定实体与属性的模式生成机制;(3) 借助语义检索与模式引导提示的丰富实体提取流程。RAVEN设计为模型无关,可根据应用需求集成不同的视觉-语言模型(VLMs)和大语言模型(LLMs)。该灵活性支持个性化搜索、内容发现和可扩展的信息检索等多样化应用,适用于海量数据的实际场景。

原文摘要 · Abstract (English)

We present RAVEN an adaptive AI agent framework designed for multimodal entity discovery and retrieval in large-scale video collections. Synthesizing information across visual, audio, and textual modalities, RAVEN autonomously processes video data to produce structured, actionable representations for downstream tasks. Key contributions include (1) a category understanding step to infer video themes and general-purpose entities, (2) a schema generation mechanism that dynamically defines domain-specific entities and attributes, and (3) a rich entity extraction process that leverages semantic retrieval and schema-guided prompting. RAVEN is designed to be model-agnostic, allowing the integration of different vision-language models (VLMs) and large language models (LLMs) based on application-specific requirements. This flexibility supports diverse applications in personalized search, content discovery, and scalable information retrieval, enabling practical applications across vast datasets.

多模态实体发现视频检索智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。