用埃洛评分优化大模型,更准地识别职场不当言论。
Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
- 用埃洛评分动态调整模型判断,提升有害内容识别能力
- 在微侵犯和仇恨言论数据集上,准确率与F1值均优于传统方法
- 适合研究职场霸凌、毒性沟通等组织安全问题的学者
大语言模型(LLMs)为组织研究提供了新机遇,但其内置的内容审核机制常导致研究人员无法分析有害内容,或产生过于保守的回应,影响结果有效性,尤其在处理职场微侵犯或仇恨言论时尤为明显。本文提出一种基于埃洛评分(Elo rating)的方法,显著提升了LLM在有害内容分析中的表现。在两个数据集上——一个聚焦微侵犯检测,另一个关注仇恨言论——该方法在准确率、精确率和F1分数等关键指标上均优于传统提示技术和常规机器学习模型。优势包括对有害内容分析的更高可靠性、更低误报率,以及对大规模数据集更强的可扩展性。该方法可支持组织应用,如识别职场骚扰、评估毒性沟通,并促进更安全、更具包容性的工作环境。
原文摘要 · Abstract (English)
Large language models (LLMs) offer promising opportunities for organizational research. However, their built-in moderation systems can create problems when researchers try to analyze harmful content, often refusing to follow certain instructions or producing overly cautious responses that undermine validity of the results. This is particularly problematic when analyzing organizational conflicts such as microaggressions or hate speech. This paper introduces an Elo rating-based method that significantly improves LLM performance for harmful content analysis In two datasets, one focused on microaggression detection and the other on hate speech, we find that our method outperforms traditional LLM prompting techniques and conventional machine learning models on key measures such as accuracy, precision, and F1 scores. Advantages include better reliability when analyzing harmful content, fewer false positives, and greater scalability for large-scale datasets. This approach supports organizational applications, including detecting workplace harassment, assessing toxic communication, and fostering safer and more inclusive work environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。