提出评估多模态主动学习陷阱的新框架,揭示现有方法易导致模态偏倚。
Mind the Gap: A Framework for Assessing Pitfalls in Multimodal Active Learning
- 用合成数据隔离缺失模态、难度差异等陷阱,实现系统性评测
- 发现模型普遍依赖单一模态,忽略其他模态,且现有策略无法缓解
- 适用于研究多模态学习与主动学习交叉问题的学者
多模态学习使神经网络能够融合异构信息源,但其主动学习面临独特挑战:模态缺失、模态难度差异以及交互结构不一致。这些在单模态场景中并不存在。尽管单模态主动学习策略的行为已得到充分研究,其在多模态条件下的表现仍不明确。本文提出一种新基准框架,通过合成数据隔离上述陷阱,实现无噪声的系统评估。利用该框架,我们对比了单模态与多模态查询策略,并在两个真实数据集上验证结果。结果显示,模型始终产生不平衡表征,主要依赖某一模态而忽视其他模态;现有查询方法未能缓解此现象,多模态策略也未持续优于单模态方案。这些发现揭示了当前主动学习方法的局限性,强调需设计显式应对上述陷阱的模态感知查询策略。代码与基准资源将公开提供。
原文摘要 · Abstract (English)
Multimodal learning enables neural networks to integrate information from heterogeneous sources, but active learning in this setting faces distinct challenges. These include missing modalities, differences in modality difficulty, and varying interaction structures. These are issues absent in the unimodal case. While the behavior of active learning strategies in unimodal settings is well characterized, their behavior under such multimodal conditions remains poorly understood. We introduce a new framework for benchmarking multimodal active learning that isolates these pitfalls using synthetic datasets, allowing systematic evaluation without confounding noise. Using this framework, we compare unimodal and multimodal query strategies and validate our findings on two real-world datasets. Our results show that models consistently develop imbalanced representations, relying primarily on one modality while neglecting others. Existing query methods do not mitigate this effect, and multimodal strategies do not consistently outperform unimodal ones. These findings highlight limitations of current active learning methods and underline the need for modality-aware query strategies that explicitly address these pitfalls. Code and benchmark resources will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。