用概念激活向量分析语音评分模型是否受母语、年龄等无关因素干扰
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
- 通过概念激活向量探测模型内部表示中是否存在母语、年龄等无关特征
- 发现模型对无关特征的敏感度取决于架构,且稀疏自编码器可提升概念可恢复性
- 强调需区分概念能否被识别与是否影响评分结果,适合关注AI公平性的研究者
自动语音评分系统在高风险场景中日益普及,用于评估第二语言学习者的口语表现,因此必须确保评分仅反映语言水平,而非母语或年龄等无关属性。基于Transformer的基座模型提升了评分准确性,但其黑箱特性使公平性和可解释性分析更具挑战。本文扩展了概念激活向量(CAVs)方法,应用于两个神经语音评分系统:基于文本的BERT评分器和基于Whisper的多模态评分器。CAVs将人类可理解的概念表示为模型激活空间中的方向,可区分概念是否被编码以及是否影响预测分数(通过梯度敏感度量化)。由于CAVs依赖线性可分性,在复杂嵌入空间中较难实现,我们进一步探究稀疏自编码器(SAEs)是否可通过稀疏潜在空间学习更清晰的概念方向并映射回激活空间。结果表明,概念可恢复性强烈依赖于所探查的表示与架构,而非概念本身;概念敏感度同样具有架构依赖性。虽然SAEs能提升概念的线性可恢复性,但会削弱原始激活空间的敏感度,尤其在低维层。这些发现强调,在审计语音评分系统的偏见时,必须区分概念可恢复性与概念影响力。
原文摘要 · Abstract (English)
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。