通过优化提示与学习集成,提升多身份视角下的毒性检测鲁棒性。
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
- 设计轻量级SVM集成模型,融合四种提示策略的输出。
- 在多种角色和模型组合下,性能超越单一提示方法和多数投票。
- 适用于需考虑社会多样性背景的主观性自然语言任务评估。
毒性检测具有固有的主观性,受不同群体的社会先验和视角影响。尽管经济学与社会科学中的“多元主义”建模旨在捕捉上下文差异,当前大语言模型的提示技术在不同角色和基础模型间表现不一。本文系统评估了面向角色的毒性检测,发现无单一提示方法在所有模型-角色组合中占优,包括我们提出的自动提示优化策略。为利用互补错误,我们探索集成四种提示变体,并提出一种轻量级元集成:基于4比特提示预测向量的SVM。结果表明,该SVM集成在多样角色中持续优于单个提示方法与传统多数投票,实现整体最优性能。本工作提供了首个针对毒性检测中角色条件提示的系统比较,并为主观性NLP任务的多元评估提供稳健方法。
原文摘要 · Abstract (English)
Toxicity detection is inherently subjective, shaped by the diverse perspectives and social priors of different demographic groups. While ``pluralistic'' modeling as used in economics and the social sciences aims to capture perspective differences across contexts, current Large Language Model (LLM) prompting techniques have different results across different personas and base models. In this work, we conduct a systematic evaluation of persona-aware toxicity detection, showing that no single prompting method, including our proposed automated prompt optimization strategy, uniformly dominates across all model-persona pairs. To exploit complementary errors, we explore ensembling four prompting variants and propose a lightweight meta-ensemble: an SVM over the 4-bit vector of prompt predictions. Our results demonstrate that the proposed SVM ensemble consistently outperforms individual prompting methods and traditional majority-voting techniques, achieving the strongest overall performance across diverse personas. This work provides one of the first systematic comparisons of persona-conditioned prompting for toxicity detection and offers a robust method for pluralistic evaluation in subjective NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。