arXiv:2604.23786cs.AIcs.LG2026-04

用可解释性提升多模态模型在心理评估中的公平性,发现透明不等于公正。

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

论文配图:FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment
图 1 · 摘自论文原文
  • 引入可解释性干预框架,通过提示词改善模型决策过程透明度
  • Phi3.5-Vision在真实场景达80.4%准确率,但存在种族偏见;Qwen2-VL性别差异更明显
  • 透明化手段能提升流程一致性,但可能加剧偏见,需兼顾准确性与公平性

近年来,多模态机器学习在心理健康评估中展现出巨大潜力。然而,视觉-语言模型(VLMs)在临床应用中因缺乏透明性和潜在偏见引发担忧。本文研究了两种不同环境下的VLM表现:实验室数据集AFAR-BSFT与自然场景数据集E-DAIC。Phi3.5-Vision在E-DAIC上达到80.4%的准确率,而Qwen2-VL仅33.9%。两个模型在AFAR-BSFT上均出现抑郁过度预测现象。尽管都存在偏见,但Qwen2-VL性别偏差更显著,Phi3.5-Vision则呈现更强的种族偏见。我们提出的可解释性干预框架效果参差:对Qwen2-VL,公平性提示实现了完美平等机会,但牺牲严重准确率;在AFAR-BSFT上,解释性干预提升流程一致性,却未保证结果公平,甚至放大种族偏见。结果表明,程序透明与结果公平之间仍存鸿沟。本文提出具体建议,强调未来公平性改进必须同时优化预测准确率、人口均等性和跨域泛化能力。

原文摘要 · Abstract (English)

In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment in clinical settings has raised concerns due to their lack of transparency and potential for bias. While previous research has explored the intersection of fairness and Explainable AI (XAI), its application to VLMs for wellbeing assessment and depression prediction remains under-explored. This work investigates VLM performance across laboratory (AFAR-BSFT) and naturalistic (E-DAIC) datasets, focusing on diagnostic reliability and demographic fairness. Performance varied substantially across environments and architectures; Phi3.5-Vision achieved 80.4% accuracy on E-DAIC, while Qwen2-VL struggled at 33.9%. Additionally, both models demonstrated a tendency to over-predict depression on AFAR-BSFT. Although bias existed across both architectures, Qwen2-VL showed higher gender disparities, while Phi-3.5-Vision exhibited more racial bias. Our XAI intervention framework yielded mixed results; fairness prompting achieved perfect equal opportunity for Qwen2-VL at a severe accuracy cost on E-DAIC. On AFAR-BSFT, explainability-based interventions improved procedural consistency but did not guarantee outcome fairness, sometimes amplifying racial bias. These results highlight a persistent gap between procedural transparency and equitable outcomes. We analyse these findings and consolidate concrete recommendations for addressing them, emphasising that future fairness interventions must jointly optimise predictive accuracy, demographic parity, and cross-domain generalisation.

多模态公平性可解释性心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。