构建情感认知视觉问答新基准,推动模型理解情绪背后原因。
InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

- 分三层标注:感知、具象理解、认知推理,覆盖情绪识别到深层意图。
- 含72.5万组问答对,30万样本用于精细评估,数据来自13.8万张筛选图像。
- 适合研究情绪理解、多模态推理与生成式模型的学者使用。
视觉情绪理解不仅要求识别情绪状态,还需解释其成因并进行高层次认知推理。现有基准多聚焦情绪识别,缺乏对具象理解与响应导向分析的支持。为此,我们提出InsightVQA,一个大规模层次化视觉问答基准,用于情绪理解与认知推理。基于6个公开来源收集的35.1万张图像,通过多阶段过滤流程筛选出13.8万张高置信度图像。每张图像在三个层级标注:感知问答(情绪与效价识别)、具象理解问答(基于视觉触发词提取,约束引导生成)、认知问答(响应意图预测与序列洞察推理)。共包含72.5万组问答对。进一步构建InsightVQA-Bench,包含3万样本的高质量评估集。为支持评估,提出InsightNet——专为多模态大模型设计的情绪调优基线。实验表明,InsightVQA对具象情绪理解与推理构成显著挑战。
原文摘要 · Abstract (English)
Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. Building from 351K images collected from six public sources, we apply a rigorous multi-stage filtering pipeline to curate 138K high-confidence images. Each image is annotated at three hierarchical levels: perception QA for emotion and valence recognition, grounded understanding QA constructed from visual trigger extraction through constraint-guided generation, and cognition QA centered on response intent prediction and sequential insight reasoning. In total, InsightVQA contains 725K QA pairs. We further present \textbf{InsightVQA-Bench}, a high-quality evaluation benchmark comprising 30K samples for fine-grained evaluation. To support evaluation, we introduce \textbf{InsightNet}, an emotion-tuned baseline for MLLMs. Results demonstrate that InsightVQA poses significant challenges for grounded emotion understanding and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。