arXiv:2603.28913cs.CL2026-03

分析大屠杀口述史中情感模型分歧,揭示敏感文本下判断不一致根源。

From Consensus to Split Decisions: ABC-Stratified Sentiment in Holocaust Oral Histories

  • 用多模型输出构建稳定性分类框架(ABC),按一致性分层
  • 整体模型一致性仅中等,关键在中立边界判断差异大
  • 适合研究历史文本情感分析的可信度与局限性

极性检测在领域迁移下挑战显著,尤其在语义复杂、结构松散的长篇叙述中,如大屠杀口述史。本文对107,305个语句、579,013个句子构成的口述史语料库,使用三种预训练Transformer极性分类器进行大规模诊断研究。通过整合模型输出,提出基于一致性的稳定性分类法(ABC),对模型间输出稳定性进行分层。报告成对一致率、Cohen kappa、Fleiss kappa及行归一化混淆矩阵,定位系统性分歧区域。作为辅助信号,采用T5-based情绪分类器对各稳定层级样本进行情绪分布比较。多模型标签三角验证与ABC分类法结合,提供一种谨慎、可操作的框架,用于刻画情感模型在敏感历史叙述中的分歧位置与机制。总体模型一致性为低至中等,主要由中立边界的判定差异驱动。

原文摘要 · Abstract (English)

Polarity detection becomes substantially more challenging under domain shift, particularly in heterogeneous, long-form narratives with complex discourse structure, such as Holocaust oral histories. This paper presents a corpus-scale diagnostic study of off-the-shelf sentiment classifiers on long-form Holocaust oral histories, using three pretrained transformer-based polarity classifiers on a corpus of 107,305 utterances and 579,013 sentences. After assembling model outputs, we introduce an agreement-based stability taxonomy (ABC) to stratify inter-model output stability. We report pairwise percent agreement, Cohen kappa, Fleiss kappa, and row-normalized confusion matrices to localize systematic disagreement. As an auxiliary descriptive signal, a T5-based emotion classifier is applied to stratified samples from each agreement stratum to compare emotion distributions across strata. The combination of multi-model label triangulation and the ABC taxonomy provides a cautious, operational framework for characterizing where and how sentiment models diverge in sensitive historical narratives. Inter-model agreement is low to moderate overall and is driven primarily by boundary decisions around neutrality.

情感分析大屠杀研究模型一致性历史文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。