arXiv:2412.08414cs.CL2024-12被引 22

用意图感知提示提升大模型对隐蔽心理操控的检测能力

Detecting Conversational Mental Manipulation with Intent-Aware Prompting

  • 基于大模型设计意图感知提示,捕捉对话中隐藏的操纵意图
  • 在MentalManip数据集上显著降低误判率,减少漏检
  • 适合心理安全监测与对话系统伦理审查场景

心理操控通过隐秘、负面的方式扭曲决策,严重损害心理健康。尽管自然语言处理领域对心理健康关注增多,但因操控手段隐蔽复杂,检测进展有限。本文提出意图感知提示(Intent-Aware Prompting, IAP),利用大语言模型深入识别对话中的操纵意图,提升检测深度。在MentalManip数据集上的实验表明,IAP优于其他先进提示策略,显著降低假阴性,能更准确识别操纵行为,同时保持正例误判极低。代码已开源。

原文摘要 · Abstract (English)

Mental manipulation severely undermines mental wellness by covertly and negatively distorting decision-making. While there is an increasing interest in mental health care within the natural language processing community, progress in tackling manipulation remains limited due to the complexity of detecting subtle, covert tactics in conversations. In this paper, we propose Intent-Aware Prompting (IAP), a novel approach for detecting mental manipulations using large language models (LLMs), providing a deeper understanding of manipulative tactics by capturing the underlying intents of participants. Experimental results on the MentalManip dataset demonstrate superior effectiveness of IAP against other advanced prompting strategies. Notably, our approach substantially reduces false negatives, helping detect more instances of mental manipulation with minimal misjudgment of positive cases. The code of this paper is available at https://github.com/Anton-Jiayuan-MA/Manip-IAP.

心理操控大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。