用精巧引用框架提升生成事实核查文章质量,二度夺冠。
UTS at CheckThat! 2026: Cite-Frame Engineering for Generated Fact-Checking Articles
- 设计双层干预机制:引用归属框架与影子验证锚点选择
- 在验证集上提升0.027分,覆盖与蕴含指标优于其他队伍
- 适合关注生成内容可信度与引用精准性的研究者
CheckThat! 2026 Task 3 要求系统生成事实核查文章,评分采用四个子指标的无权重均值(M4)。UTS提交方案在11支队伍中排名第二(M4 = 0.484)。系统为确定性草稿生成器,外接两层单杠杆干预:领域归属引用框架(HostCite)与影子验证锚点选择器(ShadowVal),仅使用 Llama-3.2:1B 作为逐条引用验证器,不用于正文生成。该架构在 WatClaimCheck 验证集上使基准草稿 M4 提升 +0.027,超越同类模型在蕴含与覆盖率上的表现。消融实验确认两项设计原则:一是评分保守性——仅认可参考文献明确支持的文本片段,模板有效,而大模型生成语句、审稿人名和原始证据均无效;二是辅助锚点信号与 Llama 判定器不匹配——所有尝试的锚点代理(交叉编码器、长度、首段位置)选出的锚点均被判为无效,应直接以判别器为准。剩余与第一名的 0.062 差距源于引用精确率/召回率(0.299 vs 0.671),符合选择性发布策略下低置信度引用被舍弃的特征。
原文摘要 · Abstract (English)
CheckThat! 2026 Task 3 asks systems to generate fact-checking articles, graded by an unweighted mean of four sub-metrics (M4). Our UTS submission placed 2nd of 11 teams (M4 = 0.484). The shipped system is a deterministic stub drafter wrapped by two single-lever interventions: a domain-attribution cite frame (HostCite) and a shadow-validated anchor picker (ShadowVal) that use Llama-3.2:1B only as a per-cite validator, never as a body-prose generator. The stack lifts M4 by +0.027 over the stub on the WatClaimCheck validation split, beats the field on entailment and coverage, and follows two design rules our ablation matrix made unambiguous. Scorer conservatism: credit only tokens the references entail - templates pay; LLM prose, reviewer names, and raw evidence all fail. Auxiliary anchor signals are miscalibrated against the Llama judge: every anchor proxy we tried (cross-encoder, length, lead position) picks anchors the judge rejects - gate on the judge itself. The remaining +0.062 gap to the winner sits on citation precision/recall (0.299 vs 0.671), consistent with a selective-emission policy that drops low-confidence cites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。