研究视觉语言模型在图像旋转下的偏差与鲁棒性问题,提出有效缓解方案。
Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models
- 通过数据增强与表征对齐提升模型对旋转的鲁棒性
- 实验显示新方法显著降低偏差放大且不损失性能
- 适合关注多模态模型公平性与鲁棒性的研究者
视觉语言模型(VLMs)和生成式图像模型在多模态任务中表现卓越,但其在输入变换下的鲁棒性与公平性仍不足。本文研究了先进多模态模型在图像旋转和分布偏移下的偏差传播与鲁棒性下降问题。分析了旋转扰动对模型预测、置信度校准及人口统计偏差模式的影响。为此,提出结合数据增强、表征对齐与模型正则化的旋转鲁棒性缓解策略。在多个数据集上的实验表明,该方法显著提升了模型鲁棒性,同时减少偏差放大,且不牺牲整体性能。研究揭示了当前多模态系统的关键局限,并提供了构建更可靠、公平AI模型的实用技术。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) and generative image models have achieved remarkable performance across multimodal tasks, yet their robustness and fairness under input transformations remain insufficiently explored. This work investigates bias propagation and robustness degradation in state-of-the-art vision-language and generative models, with a particular focus on image rotation and distributional shifts. We analyze how rotation-induced perturbations affect model predictions, confidence calibration, and demographic bias patterns. To address these issues, we propose rotation-robust mitigation strategies that combine data augmentation, representation alignment, and model-level regularization. Experimental results across multiple datasets demonstrate that the proposed methods significantly improve robustness while reducing bias amplification without sacrificing overall performance. This study highlights critical limitations of current multimodal systems and provides practical mitigation techniques for building more reliable and fair AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。