arXiv:2605.27025cs.CLcs.MM2026-05

用属性分解法提升大模型对仇恨言论的判断与人类一致

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations

论文配图:Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations
图 1 · 摘自论文原文
  • 将仇恨言论拆解为十类主观属性,逐项评估模型与人类判断的一致性
  • 显性行为属性(如侮辱、攻击)与人类标注高度相关,而评价类属性反向偏离
  • 通过加权融合属性预测,重建连续仇恨分数,性能优于直接提示方法

仇恨言论标注成本高、主观性强且标注者间分歧大,导致大规模数据集构建困难。我们系统分析了大型语言模型(LLMs)在十种理论基础明确的主观属性(如非人化、暴力、情感倾向)上与人类判断的一致性,评估了Llama 3.1和Qwen 2.5的小型与大型版本。结果发现所有模型均呈现一致分裂:行为明确维度(侮辱、羞辱、攻击-辩护)与人类标注强相关,而评价性维度(尊重、情感、仇恨言论)系统性反转。基于此,我们提出通过置信度加权的岭回归融合属性级模型预测,从Measuring Hate Speech语料库重建连续仇恨言论评分,达到最高R²=0.71,显著优于直接提示基线,表明结构化属性分解能恢复比端到端标签预测更丰富、更贴近人类判断的信号。

原文摘要 · Abstract (English)

Hate speech annotation is costly, subjective, and prone to annotator disagreement, making large-scale dataset construction challenging. We systematically analyze how well large language models (LLMs) align with human judgments across ten theoretically grounded subjective attributes, such as dehumanization, violence, and sentiment, evaluating both small and large variants of Llama 3.1 and Qwen 2.5. Our analysis reveals a consistent split across all models: behaviorally explicit dimensions (insult, humiliate, attack-defend) correlate strongly with human annotations, while evaluative dimensions (respect, sentiment, hate speech) are systematically inverted. Demographic persona conditioning reduces model confidence without improving alignment. Building on these insights, we propose combining attribute-level LLM predictions via a confidence-weighted Ridge regression to reconstruct continuous hate speech scores from the Measuring Hate Speech corpus, achieving $R^2$ of up to 0.71 and outperforming direct prompting baselines, demonstrating that structured attribute decomposition recovers a richer and more human-aligned signal than end-to-end label prediction alone.

大模型对齐仇恨言论属性分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。