动态门控让模型在少数据下更好融合视觉与语言信息
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
- 按词粒度动态调节视觉与语言信号的融合强度
- 在5个基准上表现优于或媲美主流多模态模型
- 无需监督即可发现视觉/语言侧重模式,适合小样本学习
在认知合理数据量下训练视觉-语言模型需重新思考多模态信息融合方式。针对BabyLM Challenge 2025视觉赛道约束,我们提出一种轻量级解码器架构:(1) 词粒度动态门控,自适应融合语言与视觉线索;(2) 特征调制与通道注意力,最大化有限视觉信息的利用效率;(3) 辅助对比学习目标,强化视觉定位。在五个基准(BLiMP、BLiMP Supplement、EWoK、Winoground、VQA)上的评估显示,性能优于或媲美主流多模态基线。尤为关键的是,动态门控无需显式监督即可发现可解释模式:对内容词偏好视觉线索,对功能词依赖语言线索。尽管挑战存在局限性(如全局图像嵌入造成信息瓶颈、数据集划分引发训练不稳),研究结果仍确立动态门控在高效多模态学习中的价值,兼具性能与可解释性。
原文摘要 · Abstract (English)
Training vision-language models on cognitively-plausible amounts of data requires rethinking how models integrate multimodal information. Within the constraints of the Vision track for the BabyLM Challenge 2025, we propose a lightweight decoder-based architecture with (1) token-wise dynamic gating for adaptive fusion of linguistic and visual cues, (2) feature modulation and channel attention to maximise the utility of limited visual information and (3) auxiliary contrastive objectives for visual grounding. Evaluation on five benchmarks (BLiMP, BLiMP Supplement, EWoK, Winoground and VQA) shows competitive or superior performance to multimodal baselines. More notably, our dynamic gate discovers interpretable patterns without explicit supervision, favouring visual cues for content words and linguistic cues for function words. While we identify limitations in the Challenge constraints, such as the information bottleneck created by global image embeddings and training instability from the dataset split, our findings establish dynamic gating as a powerful tool for efficient multimodal learning, offering both interpretability and performance even under severe constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。