用可解释AI提升翻译评估透明度,让机器评分更可信、更有助学习。
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
- 融合特征工程与可解释模型,只用透明特征做判断
- 在新数据集上预测准确,关键指标如BLEURT和停顿特征表现最佳
- 适合需要诊断反馈的翻译教学与自主学习场景
机器学习进步推动了自动口译质量评估的发展,但现有研究存在语言质量评估不足、数据稀缺与失衡导致建模效果不佳、且缺乏对模型预测的解释等问题。为此,我们提出一个整合特征工程、数据增强与可解释机器学习的多维建模范式。该方法通过仅使用与构念相关的透明特征,并结合Shapley值(SHAP)分析,强调可解释性而非黑箱预测。实验结果表明,在一个全新的英中连续口译数据集上,该方法表现出优异的预测性能:识别出BLEURT和CometKiwi得分是忠实度最强的预测特征,停顿相关特征对流利度预测最优,而中文特有的习语多样性指标对语言使用质量有显著影响。整体上,本研究提供了一种可扩展、可靠且透明的替代传统人工评估的方法,能为学习者提供详细诊断反馈,支持自主学习优势,这是孤立的自动化评分无法实现的。
原文摘要 · Abstract (English)
Recent advancements in machine learning have spurred growing interests in automated interpreting quality assessment. Nevertheless, existing research suffers from insufficient examination of language use quality, unsatisfactory modeling effectiveness due to data scarcity and imbalance, and a lack of efforts to explain model predictions. To address these gaps, we propose a multi-dimensional modeling framework that integrates feature engineering, data augmentation, and explainable machine learning. This approach prioritizes explainability over ``black box'' predictions by utilizing only construct-relevant, transparent features and conducting Shapley Value (SHAP) analysis. Our results demonstrate strong predictive performance on a novel English-Chinese consecutive interpreting dataset, identifying BLEURT and CometKiwi scores to be the strongest predictive features for fidelity, pause-related features for fluency, and Chinese-specific phraseological diversity metrics for language use. Overall, by placing particular emphasis on explainability, we present a scalable, reliable, and transparent alternative to traditional human evaluation, facilitating the provision of detailed diagnostic feedback for learners and supporting self-regulated learning advantages not afforded by automated scores in isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。