评测大模型识别技术术语中的有害用语能力,发现解码器模型表现更优。
An Evaluation of LLMs for Detecting Harmful Computing Terms
- 对比6种架构模型,测试其在64个技术术语场景中识别有害语言的能力。
- 解码器模型如Gemini Flash 2.0和Claude AI在上下文理解上显著优于编码器模型。
- 研究为构建更可靠的自动化包容性语言检测工具提供实证依据。
在技术领域识别有害与非包容性术语对营造包容性环境至关重要。本研究通过评估一个经筛选的技术术语数据库(每项术语配以具体使用场景),探究模型架构对有害语言检测的影响。测试了包括BERT-base-uncased、RoBERTa large-mnli、Gemini Flash 1.5 和 2.0、GPT-4、Claude AI Sonnet 3.5、T5-large、BART-large-mnli在内的多种编码器、解码器及编码器-解码器类语言模型。所有模型均采用标准化提示,针对64个术语进行有害/非包容性语言识别。结果表明,解码器模型(尤其是Gemini Flash 2.0和Claude AI)在细粒度上下文分析中表现突出;而编码器模型如BERT虽具备较强模式识别能力,但在分类置信度方面表现较弱。研究讨论了这些发现对改进自动化检测工具的启示,并强调了各类模型在促进技术领域包容性沟通中的优势与局限。
原文摘要 · Abstract (English)
Detecting harmful and non-inclusive terminology in technical contexts is critical for fostering inclusive environments in computing. This study explores the impact of model architecture on harmful language detection by evaluating a curated database of technical terms, each paired with specific use cases. We tested a range of encoder, decoder, and encoder-decoder language models, including BERT-base-uncased, RoBERTa large-mnli, Gemini Flash 1.5 and 2.0, GPT-4, Claude AI Sonnet 3.5, T5-large, and BART-large-mnli. Each model was presented with a standardized prompt to identify harmful and non-inclusive language across 64 terms. Results reveal that decoder models, particularly Gemini Flash 2.0 and Claude AI, excel in nuanced contextual analysis, while encoder models like BERT exhibit strong pattern recognition but struggle with classification certainty. We discuss the implications of these findings for improving automated detection tools and highlight model-specific strengths and limitations in fostering inclusive communication in technical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。