Vortex融合多模态信息,提升视频检索的精准与交互性。
Vortex: Multi-Modal Fusion System for Intelligent Video Retrieval

- 自适应提取关键帧,结合视觉、语言与语音生成多模态元数据
- 混合检索策略融合CLIP与SigLIP2,Reciprocal Rank Fusion提升精度
- 支持相关性反馈与分阶段时间对齐,适合智能视频问答场景
本文提出Vortex,由FocusOnFun团队为胡志明市人工智能挑战赛2025设计的多模态视频检索系统,旨在推进智能多媒体搜索与时间推理。系统集成自适应关键帧提取、基于视觉-语言与语音模型的多模态元数据生成,以及通过互斥排名融合(Reciprocal Rank Fusion)融合CLIP与SigLIP2嵌入的混合检索策略,兼顾全局与细粒度语义。为增强交互性,引入Rocchio相关性反馈与多阶段时间搜索机制,实现事件序列对齐。系统基于Milvus与Elasticsearch构建,支持可扩展索引与高效检索。在官方竞赛中,初步轮获得79.6/88(90.5%)得分,决赛轮表现优异,问题回答任务达‘卓越’水平,验证了CLIP与SigLIP2的互补优势及混合检索的有效性。该系统为未来上下文感知、交互式视频检索研究奠定坚实基础。
原文摘要 · Abstract (English)
This paper presents Vortex, the multimodal video retrieval system developed by our team, FocusOnFun, for the Ho Chi Minh City AI Challenge 2025, designed to advance intelligent multimedia search and temporal reasoning. The system integrates adaptive keyframe extraction, multimodal metadata generation from vision-language and speech models, and a hybrid retrieval strategy that fuses CLIP and SigLIP2 embeddings through Reciprocal Rank Fusion to balance global and fine-grained semantics. To enhance interactivity, Vortex incorporates Rocchio-based relevance feedback and a multi-stage temporal search mechanism for sequential event alignment. Built on Milvus and Elasticsearch, the architecture enables scalable indexing and efficient retrieval. Evaluated in the official competition, our FocusOnFun team's system achieved a score of 79.6/88 (90.5\%) in the Preliminary Round and was further evaluated in the Final Round, achieving an `Excellent' overall performance with `Outstanding' results in the question-answering (QA) task. This demonstrating the complementary strengths of CLIP and SigLIP2 and confirming the effectiveness of the hybrid retrieval approach. The system establishes a robust foundation for future research in intelligent, context-aware, and interactive video retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。