arXiv:2506.03525cs.CVcs.AI2025-06EMNLP被引 5

让视频推理模型学会分领域技能,提升跨场景理解能力。

Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning

  • 基于任务提取技能并构建多步推理链,实现领域自适应。
  • 在三个基准上均超越强基线,显著提升跨域性能。
  • 适合需要细粒度视频理解的科研与工业应用。

链式思维(CoT)推理虽提升了复杂视频理解能力,但现有方法难以适应事件检测、空间关系、情绪识别等特定领域的技能需求。为此,我们提出Video-Skill-CoT(简称Video-SKoT)框架,通过自动构建和利用技能感知的CoT监督信号,实现领域自适应视频推理。首先,从训练问题中提取领域相关推理技能,聚类形成共享技能分类体系,并为每个视频-问题对生成详细多步推理链以用于训练。其次,引入技能专用专家学习框架,每个专家模块专注于一组推理技能,使用轻量级适配器进行训练。我们在三个视频理解基准上验证了该方法的有效性,结果表明Video-SKoT持续优于多个强基线。此外,我们还深入分析了不同CoT标注流程及跨视频领域的学习技能表现。

原文摘要 · Abstract (English)

Recent advances in Chain-of-Thought (CoT) reasoning have improved complex video understanding, but existing methods often struggle to adapt to domain-specific skills (e.g., event detection, spatial relation understanding, emotion understanding) over various video content. To address this, we propose Video-Skill-CoT (a.k.a. Video-SKoT), a framework that automatically constructs and leverages skill-aware CoT supervisions for domain-adaptive video reasoning. First, we construct skill-based CoT annotations: we extract domain-relevant reasoning skills from training questions, cluster them into a shared skill taxonomy, and create detailed multi-step CoT rationale tailored to each video-question pair for training. Second, we introduce a skill-specific expert learning framework. Each expert module specializes in a subset of reasoning skills and is trained with lightweight adapters using the collected CoT supervision. We demonstrate the effectiveness of the proposed approach on three video understanding benchmarks, where Video-SKoT consistently outperforms strong baselines. We also provide in-depth analyses on comparing different CoT annotation pipelines and learned skills over multiple video domains.

视频理解链式思维领域自适应技能建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。