用大模型和搜索接口,为葡语假新闻数据集自动补充外部证据。
Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
- 用大模型提取新闻核心主张,再调用搜索接口找外部证据。
- 成功为3个葡语新闻数据集补充了外部验证信息,提升可解释性。
- 适合做葡语虚假信息检测或自动化核查系统的研究者参考。
假新闻的快速传播常常超过人工核查能力,亟需半自动化事实核查(SAFC)系统。在葡萄牙语语境下,缺乏整合外部证据的公开数据集,而现有资源多仅依赖文本内在特征进行分类。本文针对此问题,提出并应用一种方法,对葡语新闻语料库(Fake.Br、COVID19.BR、MuMiN-PT)进行外部证据增强。该方法模拟用户核查流程:使用大型语言模型(特别是Gemini 1.5 Flash)提取文本中的核心主张,并通过搜索引擎API(Google Search API、Google FactCheck Claims Search API)检索相关外部文档作为证据。同时,引入数据验证与预处理框架,包括近似重复检测,以提升原始语料质量。
原文摘要 · Abstract (English)
The accelerated dissemination of disinformation often outpaces the capacity for manual fact-checking, highlighting the urgent need for Semi-Automated Fact-Checking (SAFC) systems. Within the Portuguese language context, there is a noted scarcity of publicly available datasets that integrate external evidence, an essential component for developing robust AFC systems, as many existing resources focus solely on classification based on intrinsic text features. This dissertation addresses this gap by developing, applying, and analyzing a methodology to enrich Portuguese news corpora (Fake.Br, COVID19.BR, MuMiN-PT) with external evidence. The approach simulates a user's verification process, employing Large Language Models (LLMs, specifically Gemini 1.5 Flash) to extract the main claim from texts and search engine APIs (Google Search API, Google FactCheck Claims Search API) to retrieve relevant external documents (evidence). Additionally, a data validation and preprocessing framework, including near-duplicate detection, is introduced to enhance the quality of the base corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。