arXiv:2504.02227cs.AI2025-04AAAI被引 8

让AI回答社交问题时能看懂画面并说出理由,更像真人。

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

  • 用开放问答+视觉采样,让AI依赖图像推理
  • 在社交智能评测中准确率显著提升,且推理可信
  • 适合研究多模态社会性AI或可解释AI的学者

社交智能问答(Social-IQ)是评估模型社交智能水平的主要多模态基准。当前方法虽在封闭式选择题上取得高准确率,但严重依赖语言模态,忽视视觉上下文。且封闭选项限制了对推理路径正确性的探索。为此,我们提出视觉可解释且具象化的社会智能模型VEGAS。作为生成式多模态模型,VEGAS采用开放式回答,增强推理路径的可解释性与可评估性。通过新型帧采样策略提供更相关视觉信息,并借助通用指令微调(GIFT)提升模型对视觉帧的理解能力,目标为:一、学习基本情感社交特征的跨模态转换;二、建立多模态联合推理能力。大量实验显示,包括模态消融、开放式评估和监督式选择题测试,VEGAS均有效利用视觉信息进行推理,生成正确且可信的答案。本工作有望为Social-IQ提供新视角,推动类人社交AI发展。

原文摘要 · Abstract (English)

Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model's social intelligence level. While impressive multiple-choice question(MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entirely, dependent on language modality, overlooking visual context. Additionally, the closed-set nature further prevents the exploration of whether and to what extent the reasoning path behind selection is correct. To address these limitations, we propose the Visually Explainable and Grounded Artificial Social Intelligence (VEGAS) model. As a generative multimodal model, VEGAS leverages open-ended answering to provide explainable responses, which enhances the clarity and evaluation of reasoning paths. To enable visually grounded answering, we propose a novel sampling strategy to provide the model with more relevant visual frames. We then enhance the model's interpretation of these frames through Generalist Instruction Fine-Tuning (GIFT), which aims to: i) learn multimodal-language transformations for fundamental emotional social traits, and ii) establish multimodal joint reasoning capabilities. Extensive experiments, comprising modality ablation, open-ended assessments, and supervised MCQ evaluations, consistently show that VEGAS effectively utilizes visual information in reasoning to produce correct and also credible answers. We expect this work to of fer a new perspective on Social-IQ and advance the development of human-like social AI.

社交AI多模态可解释性视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。