arXiv:2606.09114cs.CL2026-06

通过保留语义关键锚点并校准上下文,提升中文歧视语言检测精度。

MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection

论文配图:MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection
图 1 · 摘自论文原文
  • 保留关键语义锚点,结合上下文三要素校准判断
  • 在8120样本数据集上三项指标均显著提升
  • 适合需要可解释性与轻量化部署的场景

中文歧视语言检测因有害意图常隐含且依赖上下文而具挑战性。本文提出MAAM(近视—散光锚机制),一种轻量、模型无关框架,受功能视觉模糊启发:不等同保留所有词元,而是保留与歧视相关的关键语义锚,并利用C-I-S上下文先验(上下文语气、群体身份、立场极性)进行校准。我们还构建了ChLGBT,据知是首个聚焦中文LGBT群体的歧视语言数据集,包含8,120条人工标注样本,具有显性偏见、隐性偏见和情感强度三个有序标签。在多个强编码器基线模型上,MAAM在准确率、F1、Brier分数和期望校准误差上均实现一致提升。相比前沿大模型在零样本与少样本提示下的表现,MAAM保持竞争力的同时具备更强紧凑性与稳定性。结果表明,可解释的锚点保留与上下文校准,为中文歧视语言评估提供了比模型放大更实用的替代方案。

原文摘要 · Abstract (English)

Chinese discriminatory-language detection is challenging because harmful intent is often implicit and context-dependent. We propose MAAM (Myopia--Astigmatism Anchor Mechanism), a lightweight, model-agnostic framework inspired by functional visual blur: rather than preserving every token equally, MAAM retains discrimination-relevant semantic anchors and calibrates them with C--I--S contextual priors (Contextual Tone, Group Identity, and Stance Polarity). We also introduce ChLGBT, to our knowledge the first Chinese LGBT-focused discriminatory-language dataset, with 8,120 manually annotated samples and three ordinal labels: explicit bias, implicit bias, and emotional intensity. Across strong encoder baselines, MAAM improves all three prediction dimensions, with consistent gains in accuracy, F1, Brier score, and expected calibration error. Compared with frontier LLM baselines under zero-shot and few-shot prompting protocols, MAAM remains competitive while offering stronger compactness and stability. These results suggest that interpretable anchor preservation and contextual calibration provide a practical alternative to heavier model scaling for Chinese discriminatory-language assessment.

歧视检测轻量化可解释性中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。