提出可解释可持续的文本刻板印象检测框架,提升准确率并降低碳排放。
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
- 构建多粒度刻板印象数据集EMGSD,覆盖6类群体共5.7万条标注文本。
- 基于SHAP与LIME对比生成解释置信度,确保模型输出符合人类理解。
- 在保持高精度的同时,采用碳高效ALBERT-V2模型,适合伦理与可信AI研究者。
刻板印象是关于社会群体的概括性假设,即使最先进的大语言模型在上下文学习下也难以准确识别。由于刻板印象具有主观性,其定义随文化、社会和个人视角差异而变化,因此鲁棒的可解释性至关重要。可解释模型能帮助人类用户理解并验证这些细微判断,增强信任与问责。本文提出HEARTS(可解释、可持续、鲁棒的文本刻板印象检测整体框架),在提升模型性能的同时,减少碳足迹,并提供透明可解释的推理过程。我们构建了扩展的多粒度刻板印象数据集(EMGSD),包含57,201条标注文本,涵盖六类人群,包括被代表不足的群体如LGBTQ+及地区性刻板印象。消融实验表明,在EMGSD上微调的BERT模型优于仅在单一成分上训练的模型。随后,我们对微调后的碳效率高的ALBERT-V2模型使用SHAP分析生成词元级重要性值,确保与人类理解一致,并通过比较SHAP与LIME输出计算解释置信度分数。
原文摘要 · Abstract (English)
Stereotypes are generalised assumptions about societal groups, and even state-of-the-art LLMs using in-context learning struggle to identify them accurately. Due to the subjective nature of stereotypes, where what constitutes a stereotype can vary widely depending on cultural, social, and individual perspectives, robust explainability is crucial. Explainable models ensure that these nuanced judgments can be understood and validated by human users, promoting trust and accountability. We address these challenges by introducing HEARTS (Holistic Framework for Explainable, Sustainable, and Robust Text Stereotype Detection), a framework that enhances model performance, minimises carbon footprint, and provides transparent, interpretable explanations. We establish the Expanded Multi-Grain Stereotype Dataset (EMGSD), comprising 57,201 labelled texts across six groups, including under-represented demographics like LGBTQ+ and regional stereotypes. Ablation studies confirm that BERT models fine-tuned on EMGSD outperform those trained on individual components. We then analyse a fine-tuned, carbon-efficient ALBERT-V2 model using SHAP to generate token-level importance values, ensuring alignment with human understanding, and calculate explainability confidence scores by comparing SHAP and LIME outputs...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。