用专家混合的低秩适配方法,高效生成高质量视频摘要。
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts
- 将LoRA扩展为专家混合架构,动态整合时序与空间专家。
- 在VideoXum和ActivityNet上优于现有模型,计算开销更低。
- 适合大规模视频平台需快速生成摘要的场景。
随着视频分享平台用户生成内容的爆炸式增长,高效搜索和浏览视频成为关键挑战。为帮助用户快速定位和回顾相关内容,生成简洁且信息丰富的视频摘要愈发重要。Video-llama虽能生成视频摘要,但难以有效统一建模时序与空间特征,且计算资源消耗大、耗时长。为此,我们提出MiLoRA-ViSum,通过将传统低秩适配(LoRA)扩展为复杂的专家混合范式,引入专用于视频摘要任务的双重时序-空间适应机制。该方法动态融合多个细粒度优化的LoRA专家,分别处理不同时间或空间维度。在VideoXum和ActivityNet数据集上的大量实验表明,MiLoRA-ViSum在性能上超越现有先进模型,同时显著降低计算成本。这种专家混合策略结合双重适应机制,展现出在大规模应用中兼顾效率与精度的巨大潜力。
原文摘要 · Abstract (English)
With the exponential growth of user-generated content on video-sharing platforms, the challenge of facilitating efficient searching and browsing of videos has garnered significant attention. To enhance users' ability to swiftly locate and review pertinent videos, the creation of concise and informative video summaries has become increasingly important. Video-llama is an effective tool for generating video summarization, but it cannot effectively unify and optimize the modeling of temporal and spatial features and requires a lot of computational resources and time. Therefore, we propose MiLoRA-ViSum to more efficiently capture complex temporal dynamics and spatial relationships inherent in video data and to control the number of parameters for training. By extending traditional Low-Rank Adaptation (LoRA) into a sophisticated mixture-of-experts paradigm, MiLoRA-ViSum incorporates a dual temporal-spatial adaptation mechanism tailored specifically for video summarization tasks. This approach dynamically integrates specialized LoRA experts, each fine-tuned to address distinct temporal or spatial dimensions. Extensive evaluations of the VideoXum and ActivityNet datasets demonstrate that MiLoRA-ViSum achieves the best summarization performance compared to state-of-the-art models, while maintaining significantly lower computational costs. The proposed mixture-of-experts strategy, combined with the dual adaptation mechanism, highlights the model's potential to enhance video summarization capabilities, particularly in large-scale applications requiring both efficiency and precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。