arXiv:2503.01539cs.CLcs.AI2025-03EMNLP被引 15

提出PIC提示法,提升大模型识别隐性恶意语言的能力。

Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language

  • 基于认知科学设计新提示法,引导模型进行语用推理。
  • GPT-4o等模型在隐性毒语言识别上成功率显著提升。
  • 适合需要深度语义理解的场景,如幽默、隐喻分析。

大语言模型(LLMs)的快速发展带来了伦理问题,也催生了新型毒性语言检测技术。然而,现有评估多聚焦于简单语义关联(如'he'与程序员、'she'与家庭主妇的偏见),难以应对当前更复杂的隐性毒语言形式。本研究收集了经人工标注验证的、规避网络审查的真实恶意互动数据,其具有强推理依赖特征。为评估并提升模型对这类语言的理解能力,我们提出一种新型提示方法——语用推理链(Pragmatic Inference Chain, PIC),融合认知科学与语言学成果。相比CoT、规则基线等五种基准提示,PIC显著提升了GPT-4o、Llama-3.1-70B-Instruct、DeepSeek-v2.5和DeepSeek-v3在识别隐性毒语言上的表现,并促使模型生成更清晰、连贯的推理过程,具备向幽默、隐喻等推理密集型任务迁移的潜力。

原文摘要 · Abstract (English)

The rapid development of large language models (LLMs) gives rise to ethical concerns about their performance, while opening new avenues for developing toxic language detection techniques. However, LLMs' unethical output and their capability of detecting toxicity have primarily been tested on language data that do not demand complex meaning inference, such as the biased associations of 'he' with programmer and 'she' with household. Nowadays toxic language adopts a much more creative range of implicit forms, thanks to advanced censorship. In this study, we collect authentic toxic interactions that evade online censorship and that are verified by human annotators as inference-intensive. To evaluate and improve LLMs' reasoning of the authentic implicit toxic language, we propose a new prompting method, Pragmatic Inference Chain (PIC), drawn on interdisciplinary findings from cognitive science and linguistics. The PIC prompting significantly improves the success rate of GPT-4o, Llama-3.1-70B-Instruct, DeepSeek-v2.5, and DeepSeek-v3 in identifying implicit toxic language, compared to five baseline prompts, such as CoT and rule-based baselines. In addition, it also facilitates the models to produce more explicit and coherent reasoning processes, hence can potentially be generalized to other inference-intensive tasks, e.g., understanding humour and metaphors.

大模型毒性检测语用推理隐性语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。