MERVIN统一检索越南新闻视频事件,支持多模态精准查找。
MERVIN: A Unified Framework for Multimodal Event Retrieval in Vietnamese News Videos

- 融合关键帧、字幕与摘要,构建多模态检索框架。
- 在AI Challenge HCMC 2025中获88分中的79分,决赛所有查询均正确召回。
- 支持跨模态迭代优化查询,适合越南语视频内容分析者使用。
在线视频平台的兴起推动了对高效、语义化事件检索的需求。本文提出MERVIN,一个面向越南新闻视频的统一多模态检索框架,整合关键帧、字幕和视频摘要。通过Gemini 1.5 Flash提升字幕质量,降低口音、背景音及识别错误带来的噪声。视觉特征由Perception Encoder提取,文本嵌入由越南语语言模型生成,两者均在Milvus中索引,实现高效的相似性检索。此外,基于React的交互界面支持跨模态的迭代查询优化,增强语义对齐。在越南新闻视频上的实验表明,MERVIN在AI Challenge HCMC 2025资格赛中取得88分中的79分,并在决赛中成功召回所有查询结果。
原文摘要 · Abstract (English)
The growth of online video platforms drives the need for effective, semantically grounded event retrieval. We present MERVIN, a unified multimodal framework for Vietnamese news videos that integrates keyframes, transcripts, and video summaries. Transcript quality is enhanced via Gemini 1.5 Flash, reducing noise from accents, background sounds, and recognition errors. Visual features are extracted with Perception Encoder, while a Vietnamese language model produces textual embeddings; both are indexed in Milvus for efficient similarity-based retrieval. In addition, a React-based interface enables iterative query refinement across modalities, improving semantic alignment. Experimental results on Vietnamese news videos demonstrate the effectiveness of the proposed system, with MERVIN achieving 79 out of 88 points in AI Challenge HCMC 2025 qualification phase and successfully retrieved all results for every query in the final round.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。