解决语音识别中的性别偏差问题,提升公平性与可解释性
Fair-Gate: Fairness-Aware Interpretable Risk Gating for Sex-Fair Voice Biometrics
- 通过风险外推和双分支路由机制,同时缓解性别偏见与特征纠缠
- 在VoxCeleb1上显著降低不同性别群体的验证错误率差异
- 生成可解释的路由掩码,帮助分析哪些特征影响性别判断
语音生物识别系统即使整体识别准确率高,仍可能在性别间存在性能差距。我们将其归因于两种现实机制:(i) 人口统计捷径学习,即训练过程中利用了性别与说话人身份之间的虚假相关性;(ii) 特征纠缠,即与性别相关的声学变化与身份线索重叠,无法消除而不损害身份区分能力。我们提出 Fair-Gate,一种兼顾公平性与可解释性的风险门控框架,统一处理上述两类问题。Fair-Gate 采用风险外推技术,减少不同代理性别组间的说话人分类风险差异,并引入局部互补门控机制,将中间特征分配至身份分支与性别分支。该门控生成显式的路由掩码,可供审查以理解哪些特征被用于身份或性别路径。在 VoxCeleb1 上的实验表明,Fair-Gate 改善了效用-公平性权衡,在严苛评估条件下实现了更性别公平的 ASV 性能。
原文摘要 · Abstract (English)
Voice biometric systems can exhibit sex-related performance gaps even when overall verification accuracy is strong. We attribute these gaps to two practical mechanisms: (i) demographic shortcut learning, where speaker classification training exploits spurious correlations between sex and speaker identity, and (ii) feature entanglement, where sex-linked acoustic variation overlaps with identity cues and cannot be removed without degrading speaker discrimination. We propose Fair-Gate, a fairness-aware and interpretable risk-gating framework that addresses both mechanisms in a single pipeline. Fair-Gate applies risk extrapolation to reduce variation in speaker-classification risk across proxy sex groups, and introduces a local complementary gate that routes intermediate features into an identity branch and a sex branch. The gate provides interpretability by producing an explicit routing mask that can be inspected to understand which features are allocated to identity versus sex-related pathways. Experiments on VoxCeleb1 show that Fair-Gate improves the utility--fairness trade-off, yielding more sex-fair ASV performance under challenging evaluation conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。