arXiv:2412.17505stat.MLcs.LG2024-12被引 3

揭示多模态模型中偏见的动态交互机制,发现图像能缓解文本偏见。

More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models

  • 构建仿真框架,量化文本、图像及融合模态的偏见得分。
  • 22%场景下偏见被放大,11%被缓解,67%为中性交互。
  • 图像在文本偏见主导时具稳定作用,适合公平性研究者参考。

多模态机器学习模型广泛应用于公共安全、安防和医疗等关键领域,但会继承单模态的偏见。本研究提出系统性框架分析多模态偏见的动态交互。基于包含宗教、国籍、性取向等易偏见类别的MMBias数据集,采用仿真启发式方法计算文本仅、图像仅及多模态嵌入的偏见分数。框架将偏见交互分类为放大(多模态偏见高于双单模态)、缓解(低于双单模态)和中性(介于两者之间),并进行比例分析以识别主导模态与交互动态。结果显示:当文本与图像偏见相当,22%出现放大;文本偏见占优时,11%发生缓解,凸显图像偏见的稳定作用;67%为中性,与文本偏见较高但无显著差异相关。条件概率表明,缓解场景中文本主导,中性和放大场景则存在混合贡献,揭示复杂的模态交互。研究倡导使用该可解释的系统框架,深化对跨模态偏见动态交互的理解,推动公平、可迁移的多模态模型发展。

原文摘要 · Abstract (English)

Multimodal machine learning models, such as those that combine text and image modalities, are increasingly used in critical domains including public safety, security, and healthcare. However, these systems inherit biases from their single modalities. This study proposes a systemic framework for analyzing dynamic multimodal bias interactions. Using the MMBias dataset, which encompasses categories prone to bias such as religion, nationality, and sexual orientation, this study adopts a simulation-based heuristic approach to compute bias scores for text-only, image-only, and multimodal embeddings. A framework is developed to classify bias interactions as amplification (multimodal bias exceeds both unimodal biases), mitigation (multimodal bias is lower than both), and neutrality (multimodal bias lies between unimodal biases), with proportional analyzes conducted to identify the dominant mode and dynamics in these interactions. The findings highlight that amplification (22\%) occurs when text and image biases are comparable, while mitigation (11\%) arises under the dominance of text bias, highlighting the stabilizing role of image bias. Neutral interactions (67\%) are related to a higher text bias without divergence. Conditional probabilities highlight the text's dominance in mitigation and mixed contributions in neutral and amplification cases, underscoring complex modality interplay. In doing so, the study encourages the use of this heuristic, systemic, and interpretable framework to analyze multimodal bias interactions, providing insight into how intermodal biases dynamically interact, with practical applications for multimodal modeling and transferability to context-based datasets, all essential for developing fair and equitable AI models.

多模态偏见分析公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。