arXiv:2604.20168cs.CL2026-04ACL

用大模型生成数据提升政治问答回避识别,效果显著。

Duluth at SemEval-2026 Task 6: DeBERTa with LLM-Augmented Data for Unmasking Political Question Evasions

  • 基于DeBERTa-V3,融合焦点损失与话语特征
  • 合成数据使少数类别召回率明显提升,宏F1达0.76
  • 适合研究政治话语分析与数据增强的学者

本文介绍Duluth团队在SemEval-2026任务6(CLARITY:揭示政治问答回避)中的方法。针对任务1(清晰度分类)和任务2(回避程度分类),我们使用两层标签体系对美国总统访谈中的问答对进行分类。系统基于DeBERTa-V3-base,引入焦点损失、分层学习率衰减及布尔话语特征。为缓解训练数据的类别不平衡,我们利用Gemini 3和Claude Sonnet 4.5生成合成样本。最佳配置在任务1测试集上取得0.76的宏F1,排名40支队伍中的第8位。领先系统(TeleAI)得分为0.89,平均分0.70。错误分析显示,误判主要源于模糊回应与明确回应之间的混淆,这一现象与人工标注者分歧一致。结果表明,大模型生成数据可有效提升复杂政治话语任务中少数类别的表现。

原文摘要 · Abstract (English)

This paper presents the Duluth approach to SemEval-2026 Task 6 on CLARITY: Unmasking Political Question Evasions. We address Task 1 (clarity-level classification) and Task 2 (evasion-level classification), both of which involve classifying question--answer pairs from U.S.\ presidential interviews using a two-level taxonomy of response clarity. Our system is based on DeBERTa-V3-base, extended with focal loss, layer-wise learning rate decay, and boolean discourse features. To address class imbalance in the training data, we augment minority classes using synthetic examples generated by Gemini 3 and Claude Sonnet 4.5. Our best configuration achieved a Macro F1 of 0.76 on the Task 1 evaluation set, placing 8th out of 40 teams. The top-ranked system (TeleAI) achieved 0.89, while the mean score across participants was 0.70. Error analysis reveals that the dominant source of misclassification is confusion between Ambivalent and Clear Reply responses, a pattern that mirrors disagreements among human annotators. Our findings demonstrate that LLM-based data augmentation can meaningfully improve minority-class recall on nuanced political discourse tasks.

政治话语数据增强DeBERTaLLM生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。