arXiv:2510.01502q-bio.NCcs.CV2025-10

用人类行为数据提升视频模型对社交信息的感知能力

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

  • 引入行为几何监督,约束视频嵌入的相对结构匹配人类判断
  • 在49,484次人类相似性判断上,模型性能接近噪声上限
  • 无需训练社交属性,模型自动学习情绪、支配等可解释特征

当前视频基础模型(如V-JEPA2)无法捕捉人类对动态场景中社会信息的组织方式。在多种视觉模型测试中,均无法超越基于文本描述的句子嵌入模型(MPNet)对社交视频片段的人类相似性预测能力。本文提出行为几何监督(BGS),通过约束局部与全局嵌入的成对几何关系,匹配视频间的社会关系结构。基于包含49,484个奇数挑出判断的新型人类相似性数据集,对四种ViT骨干网络(V-JEPA 2/2.1、TimeSformer、VideoMAE、CLIP)进行低秩微调。结果显示,最优微调模型V-JEPA 2.1性能接近三倍于预训练基线,并逼近噪声上限,超越最强语言嵌入基线。微调模型不仅能捕捉文本嵌入未覆盖的人类判断独特方差,还自发习得情绪(效价、唤醒度、支配力)等可解释社交情感属性,实现零样本迁移至分布外抽象社交互动数据集,并将空间注意力从场景背景转向面部、注视和交互身体区域。对照实验排除了仅靠文本蒸馏的可能。结果表明,少量人类行为数据即可引导视频模型向人类社会视觉理解靠拢。

原文摘要 · Abstract (English)

Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none were able to predict human similarity judgments to social video clips as well as a sentence embedding model of the caption text (MPNet). We show this gap in vision model performance can be closed by a compact behavioral supervisory signal. We introduce behavioral geometric supervision (BGS): a hybrid objective that constrains local and global pairwise embedding geometry to match the relational similarity structure across videos. We apply this method using a new human similarity dataset, containing 49,484 odd-one-out judgments from 250 naturalistic social video clips, and low-rank adaptation across four ViT backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, and CLIP). We find that one of the best fine-tuned models, V-JEPA 2.1, nearly triples in performance compared to the pre-trained baseline and reaches close to the noise ceiling, exceeding the strongest sentence-embedding baseline. In addition, finetuned models (i) capture unique variance in human judgments that caption-based language embeddings do not, (ii) develop interpretable social-affective attributes (valence, arousal, and dominance) despite never being trained on any of these attributes, (iii) zero-shot transfer to a separate dataset of out-of-distribution abstract social interactions, and (iv) shift spatial attention from scene context to socially informative regions (faces, gaze, and interacting bodies). A matched language-distillation control fails to reproduce these gains, ruling out caption transfer as the mechanism. Our results show how a modest amount of human behavioral data can steer video models toward human-like social visual understanding.

视频理解社会感知几何监督零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。