arXiv:2601.05545cs.CL2026-01中稿 · paper被引 3

测试大模型能否分辨有害与有争议的议论文,发现现有系统忽视伦理问题。

Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring

  • 构建包含种族、性别偏见等敏感话题的有害作文检测基准
  • 大模型和现有评分系统难以区分有害内容与合理辩论
  • 呼吁开发更关注内容伦理的自动作文评分系统

本研究针对自动作文评分(AES)系统与大语言模型(LLMs)在识别和评分有害作文方面存在的关键缺陷。尽管AES技术不断进步,当前模型常忽略文章中的伦理与道德问题,错误地给传播有害观点的作文高分。本文提出有害作文检测(HED)基准,包含涉及种族歧视、性别偏见等敏感话题的作文,用于评估多种LLMs识别与评分有害内容的能力。研究发现:(1)大语言模型仍需改进以准确区分有害与有争议的作文;(2)现有AES模型与大语言模型在评分时均未充分考虑内容的伦理维度。研究强调亟需构建对内容伦理更敏感的稳健自动评分系统。

原文摘要 · Abstract (English)

This study addresses critical gaps in Automated Essay Scoring (AES) systems and Large Language Models (LLMs) with regard to their ability to effectively identify and score harmful essays. Despite advancements in AES technology, current models often overlook ethically and morally problematic elements within essays, erroneously assigning high scores to essays that may propagate harmful opinions. In this study, we introduce the Harmful Essay Detection (HED) benchmark, which includes essays integrating sensitive topics such as racism and gender bias, to test the efficacy of various LLMs in recognizing and scoring harmful content. Our findings reveal that: (1) LLMs require further enhancement to accurately distinguish between harmful and argumentative essays, and (2) both current AES models and LLMs fail to consider the ethical dimensions of content during scoring. The study underscores the need for developing more robust AES systems that are sensitive to the ethical implications of the content they are scoring.

自动评分大模型伦理检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。