用语义感知方法更准确评估大模型在语音任务中的公平性
Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation
- 引入语义嵌入和随机效应建模,分离内容与说话人影响
- 在真实数据上显著减少虚假公平性误判,提升结果可靠性
- 适合关注语音模型公平性评估的研究者与开发者
大型音频语言模型(LALMs)在语音识别和音频问答等任务中应用日益广泛,引发对不同人口群体间公平性的担忧。在语音输入场景下,公平性评估面临语义内容变化和说话人特征等混杂因素的挑战。忽略这些因素可能导致对模型偏见的错误判断。本文提出一种语义感知的混合效应回归框架,用于评估LALMs的公平性,显式建模参考文本的句级语义嵌入作为协变量,并将说话人身份作为随机效应。值得注意的是,语义表示从待评估的LALM自身提取,实现模型视角下的语义控制。在模拟数据和真实基准上的实验表明,该方法显著减少虚假公平性发现,获得更稳健且可解释的子群性能差异估计。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation in spoken-input settings is challenging due to confounding factors, including semantic variation in spoken content and speaker-specific characteristics. Ignoring these factors can result in misleading conclusions about model bias. We propose a semantic-aware mixed-effects regression framework for fairness evaluation in LALMs that explicitly accounts for these confounders. Our approach incorporates sentence-level semantic embeddings of reference text as covariates and models speaker identity as a random effect. Notably, semantic representations are extracted from the same LALM under evaluation, enabling semantic control over variation as perceived by the model itself. Experiments on simulated data and real-world benchmarks demonstrate that the proposed approach substantially reduces spurious fairness findings and yields more robust and interpretable estimates of subgroup performance differences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。