让视频推理模型学会分领域技能,提升跨场景理解能力。
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
- 基于任务提取技能并构建多步推理链,实现领域自适应。
- 在三个基准上均超越强基线,显著提升跨域性能。
- 适合需要细粒度视频理解的科研与工业应用。
链式思维(CoT)推理虽提升了复杂视频理解能力,但现有方法难以适应事件检测、空间关系、情绪识别等特定领域的技能需求。为此,我们提出Video-Skill-CoT(简称Video-SKoT)框架,通过自动构建和利用技能感知的CoT监督信号,实现领域自适应视频推理。首先,从训练问题中提取领域相关推理技能,聚类形成共享技能分类体系,并为每个视频-问题对生成详细多步推理链以用于训练。其次,引入技能专用专家学习框架,每个专家模块专注于一组推理技能,使用轻量级适配器进行训练。我们在三个视频理解基准上验证了该方法的有效性,结果表明Video-SKoT持续优于多个强基线。此外,我们还深入分析了不同CoT标注流程及跨视频领域的学习技能表现。
原文摘要 · Abstract (English)
Recent advances in Chain-of-Thought (CoT) reasoning have improved complex video understanding, but existing methods often struggle to adapt to domain-specific skills (e.g., event detection, spatial relation understanding, emotion understanding) over various video content. To address this, we propose Video-Skill-CoT (a.k.a. Video-SKoT), a framework that automatically constructs and leverages skill-aware CoT supervisions for domain-adaptive video reasoning. First, we construct skill-based CoT annotations: we extract domain-relevant reasoning skills from training questions, cluster them into a shared skill taxonomy, and create detailed multi-step CoT rationale tailored to each video-question pair for training. Second, we introduce a skill-specific expert learning framework. Each expert module specializes in a subset of reasoning skills and is trained with lightweight adapters using the collected CoT supervision. We demonstrate the effectiveness of the proposed approach on three video understanding benchmarks, where Video-SKoT consistently outperforms strong baselines. We also provide in-depth analyses on comparing different CoT annotation pipelines and learned skills over multiple video domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。