提出融合前校准机制,让多模态信号更可信。
Before Fusion, Ask What to Keep: Contextual Calibration of Multimodal Signals

- 融合前对比各模态,生成支持与冲突信号
- 在原始特征上动态调制,提升关键信息保留率
- 适配多种模型结构,对干扰模态有鲁棒性
多模态系统常通过融合语言、声音和视觉信息提升性能,但该优势并不保证。某一输入下有用的模态在另一输入中可能成为干扰,且同一模态内的局部特征可能与其它模态证据相悖。本文研究如何在下游预测器合并前调整多模态表示。提出一个紧凑的校准模块,在摘要层面比较各模态,提取跨源支持与冲突线索,并将其转化为实例级与维度级调制信号,作用于原始模态特征而非已融合表示。该策略可抑制误导性成分、保留微弱但有用证据,并强化当前上下文支持的响应。模块为即插即用设计,可适配不同融合主干网络而不修改预测头。在涵盖情感理解、动作识别、音视频事件检测及音视频情绪分类的五个基准上,该预融合校准策略在序列与卷积融合设置下均提升性能。模态移除、合成噪声、训练动态与特征可视化分析表明,融合前校准可减少不可靠模态干扰,实现更稳定的多模态优化。
原文摘要 · Abstract (English)
Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed. A modality that is useful for one input may become distracting for another, and local feature responses within the same modality can disagree with evidence from other sources. This work investigates how to adjust multimodal representations before they are merged by a downstream predictor. We develop a compact calibration module that compares each modality with the others at the summary level, extracts cues of cross-source support and conflict, and converts these cues into instance-wise and dimension-wise modulation signals. The calibration is applied to the original modality features rather than to already fused representations, enabling the model to suppress misleading components, preserve weak but useful evidence, and emphasize responses that are better supported by the current multimodal context. The module is designed as a plug-in component and can be attached to different fusion backbones without changing their prediction heads. Across five benchmarks covering sentiment understanding, action recognition, audio-visual event detection, and audio-visual emotion classification, the proposed pre-combination calibration strategy improves performance under both sequence-based and convolutional fusion settings. Additional analyses under modality removal, synthetic corruption, training dynamics, and feature-level visualization show that calibrating signals before fusion can reduce interference from unreliable modalities and produce more stable multimodal optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。