arXiv:2606.25645cs.CRcs.AI2026-06

首次系统梳理自动化辟谣系统的32种风险,助力安全部署。

Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation

论文配图:Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation
图 1 · 摘自论文原文
  • 构建三阶段风险传播模型:风险因素→危险情境→实际危害
  • 发现32种具体风险,其中部分无法通过传统方法识别
  • 为辟谣系统安全评估提供新分析工具,适合安全研究者使用

近年来,社交网络服务(SNS)上的虚假新闻(包括错误信息和误导性信息)已成为社会问题。为应对这一问题,事实核查——即评估SNS帖子真实性——日益重要。目前的事实核查主要依赖专业机构,难以覆盖所有帖子。因此,采用自动化事实核查系统具有显著优势。然而,当前系统基于人工智能与大语言模型,存在误判风险,可能导致错误结论在社交媒体上传播,引发误导或诽谤。本文作为实现自动化事实核查系统安全应用的初步工作,提出一种风险分类框架,考虑风险从源头到危害的三阶段传播路径:风险因素、危险情境与实际损害。分析发现,自动化事实核查系统中存在32种具体风险。本文以该分类体系为分析线索(引导词),对自动化系统DEFAME进行风险评估。结果显示,利用本框架可识别出传统STRIDE方法未能覆盖的风险,验证了其有效性。

原文摘要 · Abstract (English)

In recent years, the posting of fake news including disinformation and misinformation on social networking services (SNS) has become a social problem. To combat this fake news, fact-checking that is the process of assessing the veracity of posts on SNS has become increasingly important. While fact-checking is currently performed by fact-checking organizations, it is difficult to fact-check all posts on SNS. Therefore, the use of automated fact-checking systems is effective. Recent automated fact-checking systems utilize artificial intelligence and large language models, so there are risks of incorrect judgments and posting incorrect results on social media which can lead to the spread of misinformation or to engage in defamation. In this paper, as a first step toward enabling the safe use of automated fact-checking systems, we categorize the specific risks on automated fact-checking systems. In this categorizing, we consider a three-stage risk propagation: risk factors, hazardous situations, and harm. Our analysis revealed that 32 specific risks exist in automated fact-checking systems. In this paper, we utilize the categorized risks as analytical cues (guide words) to present the risk assessment of the automated fact-checking system DEFAME. This assessment result indicates that risks that cannot be derived using STRIDE, a conventional IT security risk assessment method can be derived using our guide words.

事实核查风险评估AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。