通过跨模态一致性提升多模态模型在噪声数据下的鲁棒性。
Uncertainty-Resilient Multimodal Learning via Consistency-Guided Cross-Modal Transfer
- 用跨模态语义一致性引导特征对齐,构建共享潜在空间。
- 在情感识别任务中显著提升模型稳定性与抗噪能力。
- 适合开发高可靠性的脑机接口系统,尤其适用于低质量数据场景。
多模态学习系统常因数据噪声、标签质量差和模态异质性而面临严重不确定性,尤其在人机交互场景中,数据质量与标注一致性随用户和采集条件变化显著。本文提出基于一致性引导的跨模态迁移,以跨模态语义一致性为基础,将异构模态投影至共享潜在空间,缓解模态差异,揭示支持不确定性估计与稳定特征学习的结构关系。在此基础上,研究了提升语义鲁棒性、改善数据效率、降低噪声与不完善监督影响的策略,无需依赖大量高质量标注。在多模态情感识别基准上实验表明,该方法显著提升模型稳定性、判别能力及对噪声或不完整监督的鲁棒性。潜在空间分析进一步显示,该框架在挑战条件下仍能捕捉可靠跨模态结构。整体而言,本文整合不确定性建模、语义对齐与数据高效监督,为构建可靠自适应的脑机接口系统提供统一视角与实践启示。
原文摘要 · Abstract (English)
Multimodal learning systems often face substantial uncertainty due to noisy data, low-quality labels, and heterogeneous modality characteristics. These issues become especially critical in human-computer interaction settings, where data quality, semantic reliability, and annotation consistency vary across users and recording conditions. This thesis tackles these challenges by exploring uncertainty-resilient multimodal learning through consistency-guided cross-modal transfer. The central idea is to use cross-modal semantic consistency as a basis for robust representation learning. By projecting heterogeneous modalities into a shared latent space, the proposed framework mitigates modality gaps and uncovers structural relations that support uncertainty estimation and stable feature learning. Building on this foundation, the thesis investigates strategies to enhance semantic robustness, improve data efficiency, and reduce the impact of noise and imperfect supervision without relying on large, high-quality annotations. Experiments on multimodal affect-recognition benchmarks demonstrate that consistency-guided cross-modal transfer significantly improves model stability, discriminative ability, and robustness to noisy or incomplete supervision. Latent space analyses further show that the framework captures reliable cross-modal structure even under challenging conditions. Overall, this thesis offers a unified perspective on resilient multimodal learning by integrating uncertainty modeling, semantic alignment, and data-efficient supervision, providing practical insights for developing reliable and adaptive brain-computer interface systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。