用循环SG-MCMC与软标签学习,分解情感分类中的不确定性。
Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP

- 通过循环SG-MCMC在冻结RoBERTa上训练线性头,结合软标签目标建模标注者分布。
- 在GoEmotions数据集上,三项指标优于MC Dropout和Deep Ensemble,包括JSD、相关性与风险覆盖曲线下面积。
- 揭示硬标签校准与标注者分布契合度可独立优化,适合关注主观性评估的NLP研究者。
情感分类中的标注者分歧反映了情感概念固有的模糊性,对主观NLP中的预测器质量评估至关重要。现有工作尚未将软标签学习与贝叶斯深度学习结合,以评估包括标注者分布保真度在内的多维不确定性。本文采用循环随机梯度马尔可夫链蒙特卡洛(cSG-MCMC)在冻结的RoBERTa上训练线性头,使用软标签目标拟合经验标注者分布,并在五维评估框架下进行验证。在包含28种情感的GoEmotions基准上,所提方法在三个维度上同时优于蒙特卡洛丢弃(MC Dropout)与深度集成(Deep Ensemble):与标注者分布的杰恩申-香农散度(JSD)、每类情景下的熵不确定性和分歧的相关性(Spearman),以及选择性预测的危险-覆盖率曲线下面积(AURC)与受试者工作特征曲线下面积(AUROC)。后处理温度缩放表现出双向效应,确立硬标签校准与标注者分布契合度为独立维度,支持联合报告作为诚实评估协议。
原文摘要 · Abstract (English)
Annotator disagreement in emotion classification reflects ambiguity intrinsic to emotion concepts and is essential for predictor-quality assessment in subjective NLP. Yet no prior work integrates soft-label learning with Bayesian deep learning to evaluate uncertainty along axes including annotator-distribution fidelity. We train a linear head on a frozen RoBERTa via cyclical stochastic gradient Markov chain Monte Carlo (cSG-MCMC), targeting the empirical annotator distribution with a soft-label objective under a five-axis evaluation. On the 28-emotion GoEmotions benchmark, the proposed method outperforms Monte Carlo Dropout and Deep Ensemble simultaneously on three axes -- Jensen-Shannon divergence (JSD) to the annotator distribution, Spearman correlation between per-emotion aleatoric uncertainty and disagreement, and selective-prediction Area Under the Risk-Coverage Curve (AURC) and Area Under the ROC Curve (AUROC) -- showing independent axes are jointly attainable from one posterior. Post-hoc temperature scaling exhibits a bidirectional effect, establishing hard-label calibration and annotator-JSD as independent dimensions and motivating joint reporting as an honest protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。