解决视觉语言模型社交能力退化问题,提升多任务协同表现
SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- 冻结视觉编码器,仅训练轻量连接模块以缓解社交信息丢失
- 在五个社交任务上实现正向迁移,性能接近专用模型
- 揭示预训练导致社交表征能力下降,适合研究社会智能的学者
从视觉线索理解社会互动是构建具备社会智能AI的核心挑战。尽管强大的预训练视觉语言模型(VLMs)展现出卓越的泛化能力,却在同时学习多个社会感知任务时表现不佳,常出现负向迁移。我们发现这种负向迁移源于一种关键问题——“社交退化”,即VLM的通用视觉-语言预训练过程损害了视觉编码器对细微社交信息的表征能力。通过线性表示探测和梯度冲突分析两种视角深入研究,发现可解码性显著下降是主要因素。为此,我们提出SocialFusion框架,通过在冻结的视觉编码器与语言模型间建立最小连接,实现统一建模。相比现有VLMs,该方法在全部五个社交任务上均实现正向迁移,利用任务间协同提升整体性能,并在多个基准上达到与特定任务最优模型相当的水平。结果表明,当前VLM预训练策略可能不利于获得通用社会认知能力,亟需更注重社会性的训练范式。
原文摘要 · Abstract (English)
Understanding social interactions from visual cues is a fundamental challenge for a socially competent AI. While powerful pre-trained vision-language models (VLMs) have shown remarkable general capabilities, they surprisingly struggle to unify and learn multiple social perception tasks simultaneously, often exhibiting negative transfer. We identify that this negative transfer stems from a critical issue we term "social degradation," whereby the general visual-linguistic pre-training process of VLMs impairs the visual encoder's ability to represent nuanced social information. We investigate this behavior further under two lenses: decodability through linear representation probing and compatibility through gradient conflict analysis, revealing that both play a role in the degradation, especially the former, which is significantly compromised in the VLM pre-training process. To address these issues, we propose SocialFusion, a unified framework that learns a minimal connection between a frozen visual encoder and a language model. Compared with existing VLMs, it exhibits positive transfer across all five social tasks, leveraging synergies between them to enhance overall performance and achieves comparable performance to task-specific state-of-the-art models on various benchmarks. Our findings suggest that current VLM pre-training strategies may be detrimental to acquiring general social competence and highlight the need for more socially-aware training paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。