arXiv:2503.05605cs.IRcs.CY2025-03

自动识别维基百科流中谣言并实时解释,提升内容审核效率。

Identification and explanation of disinformation in wiki data streams

  • 基于流式处理与特征工程,实现在线内容质量验证
  • 在两个数据集上各项指标均达90%以上准确率
  • 提供公众可读的解释面板,适合编辑与普通用户使用

社交媒体作为新闻来源日益普及,其内容未经验证,易传播虚假信息,影响读者判断力。本文针对维基百科页面内容流快速增长的问题,提出一种可扩展的自动数据质量验证方案,涵盖流式数据处理、特征工程与选择、流式分类及预测结果的实时解释。系统设计了面向大众的可解释性仪表板,帮助无专业背景的用户理解模型判断。在两个数据集上的实验结果显示,所有评估指标均达到约90%的性能,优于现有研究。该系统可显著减少编辑发现虚假信息所需的时间与精力。

原文摘要 · Abstract (English)

Social media platforms, increasingly used as news sources for varied data analytics, have transformed how information is generated and disseminated. However, the unverified nature of this content raises concerns about trustworthiness and accuracy, potentially negatively impacting readers' critical judgment due to disinformation. This work aims to contribute to the automatic data quality validation field, addressing the rapid growth of online content on wiki pages. Our scalable solution includes stream-based data processing with feature engineering, feature analysis and selection, stream-based classification, and real-time explanation of prediction outcomes. The explainability dashboard is designed for the general public, who may need more specialized knowledge to interpret the model's prediction. Experimental results on two datasets attain approximately 90 % values across all evaluation metrics, demonstrating robust and competitive performance compared to works in the literature. In summary, the system assists editors by reducing their effort and time in detecting disinformation.

信息检测流式处理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。