用大模型识别反犹事件,发现效果有潜力但需改进。
Evaluating Large Language Models for Antisemitic Incident Classification

- 通过定义和示例优化提示词,提升模型对反犹事件的分类能力。
- GPT-4o在多数据集上表现优于Llama-3.2-3B-Instruct,但仍需改进。
- 适合政策制定者、公益组织及AI开发者用于早期预警系统建设。
应对社会中的仇恨与暴力,亟需从公开报告中及时检测仇恨事件,但自动化识别仍研究不足。本文提出仇恨事件检测任务,考察大语言模型(LLMs)对反犹事件报告进行细粒度分类的能力。我们评估了OpenAI的GPT-4o与Meta的Llama-3.2-3B-Instruct在多个专家标注的数据集上的表现,这些数据来自新闻报道、民间组织报告及官方记录。结果显示,尽管大模型具有潜力,尤其是GPT-4o,但仍有显著提升空间。在提示词中加入清晰术语定义可提升对修辞类事件(如经典反犹隐喻)的识别;提供上下文示例则更有利于行动类事件(如身体袭击)的判断。对高校校报的案例研究显示,模型能有效挖掘真实世界事件,支持早期监测与干预。整体表明,当前AI在识别复杂仇恨行为方面存在机遇与关键短板,亟需AI开发者、政策制定者与民间社会协同推进模型设计、严谨评估与政策框架建设。
原文摘要 · Abstract (English)
Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI's GPT-4o and Meta's Llama-3.2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault). A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI's ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。