arXiv:2503.03663cs.CV2025-03CVPR被引 68

LION-FS让视频助手实时回应更准更细,兼顾速度与效果。

LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

  • 分快慢两路:快路判断是否该回,慢路精细提取关键帧特征。
  • 在真实视频上实现毫秒级响应,准确率超现有方法。
  • 适合做在线视频对话的智能助理,尤其擅长复杂动作理解。

第一人称视频助手有望通过在线视频对话提升日常生活体验。然而,现有在线视频助手常为追求实时性而降低效能,采用低帧率视频和粗粒度视觉特征。为突破效能与效率的权衡,本文提出“快慢结合”的视频语言思考者 LION-FS,实现实时、主动、时序精准且上下文精确的响应。LION-FS 采用两阶段优化策略:1)快路径:基于路由的响应判定,逐帧判断是否需立即响应;通过令牌聚合路由动态融合时空特征,不增加令牌数,同时使用令牌丢弃路由消除冗余特征;2)慢路径:多粒度关键帧增强,在响应生成中优化关键帧;通过多粒度池化提取细粒度空间特征及人-环境交互特征,并融入精心设计的多模态思维模板,以指导更精准的响应生成。在在线视频任务上的综合评估表明,LION-FS 达到当前最优的效能与效率水平。

原文摘要 · Abstract (English)

First-person video assistants are highly anticipated to enhance our daily lives through online video dialogue. However, existing online video assistants often sacrifice assistant efficacy for real-time efficiency by processing low-frame-rate videos with coarse-grained visual features.To overcome the trade-off between efficacy and efficiency, we propose "Fast & Slow Video-Language Thinker" as an onLIne videO assistaNt, LION-FS, achieving real-time, proactive, temporally accurate, and contextually precise responses. LION-FS adopts a two-stage optimization strategy: 1)Fast Path: Routing-Based Response Determination evaluates frame-by-frame whether an immediate response is necessary. To enhance response determination accuracy and handle higher frame-rate inputs efficiently, we employ Token Aggregation Routing to dynamically fuse spatiotemporal features without increasing token numbers, while utilizing Token Dropping Routing to eliminate redundant features. 2)Slow Path: Multi-granularity Keyframe Augmentation optimizes keyframes during response generation. To provide comprehensive and detailed responses beyond atomic actions constrained by training data, fine-grained spatial features and human-environment interaction features are extracted through multi-granular pooling. These features are further integrated into a meticulously designed multimodal Thinking Template to guide more precise response generation. Comprehensive evaluations on online video tasks demonstrate that LION-FS achieves state-of-the-art efficacy and efficiency.

视频助手多模态实时推理关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。