arXiv:2512.09563cs.CLcs.AI2025-12中稿 · CCL 2025

用三阶段框架提升中文仇恨言论细粒度识别能力

System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection

  • 设计上下文感知提示,引导大模型提取隐含仇恨模式
  • 融合任务特征微调,显著提升对中文网络用语的识别率
  • 多模型合并增强对新变种仇恨言论的鲁棒性

中文社交媒体上仇恨言论泛滥,传统系统难以应对依赖语境的修辞策略和不断演变的网络用语。为此,我们提出一种三阶段基于大模型的框架:提示工程、监督微调与大模型合并。首先,设计上下文感知提示,引导大模型提取隐含仇恨模式;其次,在监督微调中融入任务特定特征,增强领域适应性;最后,通过合并微调后的多个大模型,提升对分布外情况的鲁棒性。在STATE-ToxiCN基准上的评估验证了该框架的有效性,其在细粒度仇恨言论检测任务中表现优于基线方法。

原文摘要 · Abstract (English)

The proliferation of hate speech on Chinese social media poses urgent societal risks, yet traditional systems struggle to decode context-dependent rhetorical strategies and evolving slang. To bridge this gap, we propose a novel three-stage LLM-based framework: Prompt Engineering, Supervised Fine-tuning, and LLM Merging. First, context-aware prompts are designed to guide LLMs in extracting implicit hate patterns. Next, task-specific features are integrated during supervised fine-tuning to enhance domain adaptation. Finally, merging fine-tuned LLMs improves robustness against out-of-distribution cases. Evaluations on the STATE-ToxiCN benchmark validate the framework's effectiveness, demonstrating superior performance over baseline methods in detecting fine-grained hate speech.

仇恨言论检测大模型融合中文NLP提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。