arXiv:2511.07988cs.AI2025-11被引 3

用脑活动数据优化多模态模型,提升对情景喜剧讽刺的识别能力。

The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends

  • 基于脑功能区活动数据微调音视频模型,增强社会认知理解。
  • 在《老友记》观看中,脑区对齐度显著提升,讽刺识别准确率提高。
  • 适合关注脑机接口与多模态智能融合的研究者。

近期研究表明,通过脑活动预测目标对模型进行脑调优(brain-tuning),可提升模型与大脑活动的一致性,并改善下游语义与音频任务表现。本研究将该方法扩展至多模态音视频模型,聚焦于社会认知关键脑区——上颞沟(STS),让受试者观看《老友记》时进行调优。结果显示,模型在STS及其邻近区域的脑活动对齐度显著提升,同时在训练数据相关的讽刺识别任务中性能明显改善。研究证明,将脑调优应用于多模态场景,能有效增强模型的社会认知能力。

原文摘要 · Abstract (English)

Recent studies on audio models show brain-tuning - fine-tuning models to better predict corresponding fMRI activity - improves brain alignment and increases performance on downstream semantic and audio tasks. We extend this approach to a multimodal audio-video model to enhance social cognition, targeting the Superior Temporal Sulcus (STS), a key region for social processing, while subjects watch Friends. We find significant increases in brain alignment to the STS and an adjacent ROI, as well as improvements to a social cognition task related to the training data - sarcasm detection in sitcoms. In summary, our study extends brain-tuning to the multi-modal domain, demonstrating improvements to a downstream task after tuning to a relevant functional region.

脑调优多模态社会认知讽刺识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。