arXiv:2605.05640cs.CV2026-05

让AI主动找长视频中的情绪片段,回应模糊提问。

AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries

论文配图:AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
图 1 · 摘自论文原文
  • 设计智能体分步查找情绪片段并验证
  • 在长视频中准确定位情绪时刻并生成理由
  • 适合需要理解复杂视频情感的交互场景

现有情感理解研究多集中于图像、音频或预剪辑视频片段的情绪识别,其情感证据已预先给出。这种被动、片段中心的设定无法反映真实场景:用户常与长视频交互,并通过自然语言提问表达需求。本文提出新任务「模糊查询驱动的视频情感理解(VQAU)」,要求模型在长视频中定位情感时刻、预测情绪类别并生成基于证据的理由。为此构建了统一评估框架VQAU-Bench,包含长视频、模糊情感查询、时间片段标注、情绪标签和理由解释。为应对多步推理挑战,提出AffectSeek框架,通过角色专精推理与跨阶段验证,将模糊用户意图逐步对齐长视频证据。实验表明,现有模型在该任务上表现不佳,而AffectSeek提供了简单有效的智能体式解决方案。

原文摘要 · Abstract (English)

Existing affective understanding studies have mainly focused on recognizing emotions from images, audio signals, or pre-cliped video clips, where the affective evidence is already given. This passive and clip-centered setting does not fully reflect real-world scenarios, in which users often interact with long videos and express their needs through natural-language queries. In this paper, we study \textbf{Vague-Query-driven video Affective Understanding (VQAU)}, a new task that requires models to localize affective moments in long videos, predict their emotion categories, and generate evidence-grounded rationales under vague user queries. To support this task, we construct \textbf{VQAU-Bench}, a benchmark that integrates long videos, vague affective queries, temporal clip annotations, emotion labels, and rationale explanations into a unified evaluation framework. VQAU-Bench enables systematic assessment of semantic-temporal-affective alignment, affective moment localization, emotion classification, and rationale generation. To address the multi-step reasoning challenges of VQAU, we further propose \textbf{AffectSeek}, an agentic framework that actively seeks, verifies, and explains affective moments in long videos. AffectSeek decomposes VQAU into intent interpretation, candidate localization, clip verification, emotion reasoning, and rationale generation, and progressively aligns vague user intent with long-video evidence through role-specialized reasoning and cross-stage verification. Experiments show that VQAU remains challenging for existing affective recognition models and single-step vision-language models, while AffectSeek provides a simple yet effective framework for agentic long-video affective understanding.

情感理解长视频智能体自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。