arXiv:2508.06360cs.CL2025-08中稿 · RANLP 2025被引 3

用攻击性检测增强提示,提升大模型识别网络霸凌的能力。

Cyberbullying Detection via Aggression-Enhanced Prompting

  • 将攻击性检测作为辅助任务,嵌入提示中提供上下文增强。
  • 在五个攻击性数据集和一个霸凌数据集上表现优于传统微调方法。
  • 适合关注社交媒体安全、大模型泛化能力的研究者。

由于网络霸凌表达方式微妙多变,其检测仍是重大挑战。本研究探究在统一训练框架中引入攻击性检测作为辅助任务,能否提升大语言模型(LLMs)在霸凌检测中的泛化能力与性能。实验在五个攻击性数据集和一个霸凌数据集上使用指令微调的LLM进行,评估了零样本、少样本、独立LoRA微调及多任务学习(MTL)等多种策略。鉴于MTL结果不一致,提出一种增强提示管道方法,将攻击性预测结果融入霸凌检测提示以提供上下文增益。初步结果表明,该方法在所有测试场景下均优于标准LoRA微调,说明基于攻击性的上下文显著提升了霸凌检测效果。研究凸显了攻击性检测等辅助任务在提升社交网络安全应用中大模型泛化能力方面的潜力。

原文摘要 · Abstract (English)

Detecting cyberbullying on social media remains a critical challenge due to its subtle and varied expressions. This study investigates whether integrating aggression detection as an auxiliary task within a unified training framework can enhance the generalisation and performance of large language models (LLMs) in cyberbullying detection. Experiments are conducted on five aggression datasets and one cyberbullying dataset using instruction-tuned LLMs. We evaluated multiple strategies: zero-shot, few-shot, independent LoRA fine-tuning, and multi-task learning (MTL). Given the inconsistent results of MTL, we propose an enriched prompt pipeline approach in which aggression predictions are embedded into cyberbullying detection prompts to provide contextual augmentation. Preliminary results show that the enriched prompt pipeline consistently outperforms standard LoRA fine-tuning, indicating that aggression-informed context significantly boosts cyberbullying detection. This study highlights the potential of auxiliary tasks, such as aggression detection, to improve the generalisation of LLMs for safety-critical applications on social networks.

网络霸凌大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。