arXiv:2603.14027cs.CL2026-03ACL被引 13

评测政治问答中的模糊回避行为,揭示大模型在识别策略性回避上的局限。

SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions

  • 构建两级分类任务:回答清晰度与回避策略细粒度识别
  • 清晰回答分类最优模型达0.89宏F1,回避策略仅0.68
  • 利用层级分类框架和提示工程是最佳策略,独立处理效果差

政治发言者常以看似回应的方式回避问题,但此类策略性回避在自然语言处理中仍研究不足。我们提出SemEval-2026 Task 6 CLARITY,一项关于政治问答回避的共享任务,包含两个子任务:(i) 回答清晰度分类(清晰回应、模糊、清晰不回应),(ii) 回避策略细粒度分类(九种策略)。数据集基于美国总统访谈构建,采用专家定义的清晰度与回避分类体系。共有124支队伍注册,提交了946次有效清晰度分类运行和539次回避分类运行。结果表明两任务难度差异显著:最优系统在清晰度分类上达到0.89宏F1,远超最强基线;而回避分类最佳系统仅达0.68宏F1,与最强基线持平。整体来看,大语言模型提示与分类体系的层级利用最有效,顶尖系统始终优于独立处理两任务的方法。CLARITY将政治话语中的策略性模糊确立为计算话语分析的挑战性基准,凸显建模政治语言模糊性的困难。

原文摘要 · Abstract (English)

Political speakers often avoid answering questions directly while maintaining the appearance of responsiveness. Despite its importance for public discourse, such strategic evasion remains underexplored in Natural Language Processing. We introduce SemEval-2026 Task 6, CLARITY, a shared task on political question evasion consisting of two subtasks: (i) clarity-level classification into Clear Reply, Ambivalent, and Clear Non-Reply, and (ii) evasion-level classification into nine fine-grained evasion strategies. The benchmark is constructed from U.S. presidential interviews and follows an expert-grounded taxonomy of response clarity and evasion. The task attracted 124 registered teams, who submitted 946 valid runs for clarity-level classification and 539 for evasion-level classification. Results show a substantial gap in difficulty between the two subtasks: the best system achieved 0.89 macro-F1 on clarity classification, surpassing the strongest baseline by a large margin, while the top evasion-level system reached 0.68 macro-F1, matching the best baseline. Overall, large language model prompting and hierarchical exploitation of the taxonomy emerged as the most effective strategies, with top systems consistently outperforming those that treated the two subtasks independently. CLARITY establishes political response evasion as a challenging benchmark for computational discourse analysis and highlights the difficulty of modeling strategic ambiguity in political language.

政治话语回避识别分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。