发现大模型在孟加拉语标注中因指令导致标签坍缩,误判率超七成。
MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media

- 用四大模型分批标注,共享验证集诊断偏差
- 79%仇恨内容、75%讽刺内容被模型漏检
- 标签高一致实为假象,适合关注低资源语言公平性研究者
通过大型语言模型(LLMs)实现标注自动化是扩展NLP数据集的核心方法;然而,低资源语言中封闭式指令下LLM行为尚未得到充分研究。我们提出MultiSoc-4D,一个包含58,000+条来自六个来源的孟加拉语社交媒体评论的数据集,涵盖类别、情感、仇恨言论和讽刺四维标注。采用结构化流程,ChatGPT、Gemini、Claude与Grok分别标注不同分区,共享20%公共验证集,系统诊断LLM行为。我们发现一种普遍现象——“指令诱导标签坍缩”,即模型系统性倾向使用备用标签(Other、Neutral、No),导致高一致性但显著低估少数类别。例如,模型对仇恨内容和讽刺内容的漏检率分别达79%和75%,相较人类校准参考。进一步证明该现象为“标签一致假象”,讽刺检测的Fleiss' Kappa接近零(κ≈−0.001)。在40+种模型上验证了此标注偏差在训练流程中的传播,不受架构差异影响。我们发布MultiSoc-4D作为孟加拉语NLP中注释偏差的诊断基准。
原文摘要 · Abstract (English)
Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set instructions in low-resource languages has not been well studied. We present MultiSoc-4D, a Bengali social media dataset benchmark, which contains 58K+ social media comments from six sources annotated along four dimensions: category, sentiment, hate speech, and sarcasm. By employing a structured pipeline where ChatGPT, Gemini, Claude, and Grok individually annotate separate partitions, while sharing a common validation set of 20%, we diagnose LLM behavior systematically. We discover a prevalent phenomenon called "instruction-induced label collapse", wherein LLMs show a systematic preference towards fallback labels (Other, Neutral, No), leading to high agreement rates but under-detection of minority categories. For example, we find that LLMs failed to detect 79% and 75% of instances with hateful and sarcastic content compared to a human-calibrated reference. Furthermore, we prove that it represents a "label agreement illusion", statistically validated via almost null Fleiss' Kappa ($κ\approx -0.001$) on sarcasm detection. Across 40+ LLMs, we benchmark this annotation bias propagation within the training pipeline, regardless of architectural differences. We release MultiSoc-4D as a diagnostic benchmark for annotation biases in Bengali NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。