arXiv:2505.15255cs.CL2025-05中稿 · WWW 2026被引 3

用数据增强与蒸馏提升大模型识别社交操纵能力

Boosting Large Language Models for Mental Manipulation Detection via Data Augmentation and Distillation

  • 无标注数据增强结合语用提示,生成高质量训练样本
  • 在真实对话数据上实现准确率提升14.0%、宏平均F1提升27.3%
  • 适合安全检测、内容审核方向的研究与应用

社交媒体中的心理操纵行为对个体心理健康和网络互动完整性构成隐蔽而严重的威胁。由于标注数据稀缺、行为隐匿且多轮交互复杂,检测难度高,且缺乏真实世界数据集。为此,我们提出MentalMAD框架,通过三个核心组件提升大语言模型的检测能力:EvoSA——一种无需标注的数据增强方法,结合进化操作与语用感知提示;教师模型生成的互补任务监督;以及分阶段的互补-收敛蒸馏策略,将操纵知识有效迁移至学生模型。我们构建了包含5000条真实来源对话的ReaMent数据集。大量实验表明,MentalMAD相比最强基线,准确率提升14.0%,宏平均F1提升27.3%,加权F1提升15.1%。代码与数据集已公开于https://github.com/Yuansheng-Gao/MentalMAD。

原文摘要 · Abstract (English)

Mental manipulation on social media poses a covert yet serious threat to individuals' psychological well-being and the integrity of online interactions. Detecting such behavior is challenging due to the difficult-to-annotate training data, its highly covert and multi-turn nature, and the lack of real-world datasets. To address these challenges, we propose MentalMAD, a framework that enhances large language models for mental manipulation detection. Our approach consists of three key components: EvoSA, an annotation-free data augmentation method that combines evolutionary operations with speech-act-aware prompting; teacher-model-generated complementary-task supervision; and Complementary-Convergent Distillation, a phase-wise strategy for transferring manipulation-specific knowledge to student models. We then constructed the ReaMent dataset, comprising 5,000 real-world-sourced dialogues. Extensive experiments show that MentalMAD improves accuracy by 14.0%, macro-F1 by 27.3%, and weighted F1 by 15.1% over the strongest baseline. The code and the dataset are publicly available at https://github.com/Yuansheng-Gao/MentalMAD.

心理操纵大模型数据增强内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。