arXiv:2603.12453cs.CL2026-03ACL被引 1

用双模型集成与复杂度门控提升政治回应模糊性检测准确率

CSE-UOI at SemEval-2026 Task 6: A Two-Stage Heterogeneous Ensemble with Deliberative Complexity Gating for Political Evasion Detection

  • 采用双大模型自一致性与加权投票的异构集成方法
  • 通过响应长度代理实现模糊性判别,达0.85宏平均F1
  • 适合关注政治话语分析与模型推理机制的研究者

本文介绍我们在SemEval-2026任务6中的系统,该任务将政治访谈回应的清晰度分为三类:明确回答、模棱两可和明确回避。我们提出一种基于自一致性(SC)和加权投票的异构双大语言模型(LLM)集成方法,并引入一种新型后处理修正机制——思辨复杂度门控(DCG)。该机制利用跨模型行为信号,发现响应长度代理与样本模糊性存在强相关性。为进一步探究提升模糊性检测的机制,我们评估了多智能体辩论作为增强思辨能力的替代策略。与依赖跨模型行为信号动态调控推理的DCG不同,辩论通过增加智能体数量但不提升模型多样性来增强推理。我们的方案在测试集上取得0.85的宏平均F1分数,获得第三名,与第二名报告得分并列。

原文摘要 · Abstract (English)

This paper describes our system for SemEval-2026 Task 6, which classifies clarity of responses in political interviews into three categories: Clear Reply, Ambivalent, and Clear Non-Reply. We propose a heterogeneous dual large language model (LLM) ensemble via self-consistency (SC) and weighted voting, and a novel post-hoc correction mechanism, Deliberative Complexity Gating (DCG). This mechanism uses cross-model behavioral signals and exploits the finding that an LLM response-length proxy correlates strongly with sample ambiguity. To further examine mechanisms for improving ambiguity detection, we evaluated multi-agent debate as an alternative strategy for increasing deliberative capacity. Unlike DCG, which adaptively gates reasoning using cross-model behavioral signals, debate increases agent count without increasing model diversity. Our solution achieved a Macro-F1 score of 0.85 on the evaluation set, securing 3rd place and tied with the second-best reported score.

政治话语模糊检测大模型集成推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。