arXiv:2509.06883cs.CLcs.AI2025-09

对比微调与提示法,发现高分不等于高质量的谣言检测文本提取

UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction

  • 对比不同大模型的微调与提示策略,测试其在社交媒体中提取待验证观点的效果
  • 最佳METEOR得分来自FLAN-T5模型的微调,但部分非最优方法也能提取更高质量语句
  • 对内容质量与指标分数不一致现象提出警示,适合关注事实核查系统设计的研究者

我们参与了CheckThat! 2025任务2(英文),探索多种提示方法与上下文学习策略,包括少样本提示和不同大语言模型家族的微调,旨在从社交媒体文本中提取值得验证的声明。最佳METEOR得分由微调FLAN-T5模型获得。然而,我们观察到,即使某些方法的METEOR得分较低,仍可能提取出质量更高的声明。

原文摘要 · Abstract (English)

We participate in CheckThat! Task 2 English and explore various methods of prompting and in-context learning, including few-shot prompting and fine-tuning with different LLM families, with the goal of extracting check-worthy claims from social media passages. Our best METEOR score is achieved by fine-tuning a FLAN-T5 model. However, we observe that higher-quality claims can sometimes be extracted using other methods, even when their METEOR scores are lower.

信息提取大模型应用事实核查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。