arXiv:2511.12255cs.CV2025-11

Fusionista2.0加速视频检索,75%提速同时提升准确率与用户体验。

Fusionista2.0: Efficiency Retrieval System for Large-Scale Datasets

  • 全模块重设计:用ffmpeg、faster-whisper等工具实现高效预处理与识别
  • 检索速度最高降低75%,准确率和用户满意度双提升
  • 新界面更易用,非专业人士也能快速查到所需视频内容

Video Browser Showdown (VBS) 要求系统在严格时间限制下提供精准结果。为应对这一挑战,我们提出Fusionista2.0,一个面向大规模数据集的高效视频检索系统。所有核心模块均重新优化:预处理采用ffmpeg实现快速关键帧提取,文字识别使用Vintern-1B-v3.5进行鲁棒的多语言文本识别,语音识别则采用faster-whisper实现实时转录。问答环节采用轻量级视觉语言模型,快速响应且无需大型模型的高昂开销。此外,Fusionista2.0引入全新用户界面,显著提升响应速度、可访问性与工作流效率,使非专家用户也能迅速获取相关视频内容。评估显示,检索时间最多减少75%,准确率与用户满意度同步上升,验证了该系统在大规模视频搜索中的竞争力与友好性。

原文摘要 · Abstract (English)

The Video Browser Showdown (VBS) challenges systems to deliver accurate results under strict time constraints. To meet this demand, we present Fusionista2.0, a streamlined video retrieval system optimized for speed and usability. All core modules were re-engineered for efficiency: preprocessing now relies on ffmpeg for fast keyframe extraction, optical character recognition uses Vintern-1B-v3.5 for robust multilingual text recognition, and automatic speech recognition employs faster-whisper for real-time transcription. For question answering, lightweight vision-language models provide quick responses without the heavy cost of large models. Beyond these technical upgrades, Fusionista2.0 introduces a redesigned user interface with improved responsiveness, accessibility, and workflow efficiency, enabling even non-expert users to retrieve relevant content rapidly. Evaluations demonstrate that retrieval time was reduced by up to 75% while accuracy and user satisfaction both increased, confirming Fusionista2.0 as a competitive and user-friendly system for large-scale video search.

视频检索效率优化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。