构建开放可信度数据集,助力浏览器实时识别虚假信息
CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-Bunking of Online Misinformation
- 融合四个公开信号源构建域名可信度评分体系
- 覆盖2672个域名,评分0.0~1.0,支持精准分类
- 专为隐私保护的客户端部署设计,适合安全工具开发
本文提出CRED-1,一个开源、可复现的领域级可信度数据集,整合两个公开授权来源(OpenSources.co与Iffy.news)及四种计算增强信号:域名年龄(WHOIS/RDAP)、网页流行度(Tranco Top-1M)、事实核查频率(Google Fact Check Tools API)和威胁情报(Google Safe Browsing API)。该数据集涵盖2,672个域名,分为虚假、不可靠、混合、阴谋论和讽刺五类,每个域名获得0.0至1.0之间的综合可信度评分。CRED-1专为私密保护的浏览器扩展设计,支持在内容传输阶段实现客户端预反制。整个流程使用Python标准库实现,完全可从公开资源复现。数据集与代码以CC BY 4.0许可发布,并存档于Zenodo。
原文摘要 · Abstract (English)
This article presents CRED-1, an open, reproducible domain-level credibility dataset combining two openly-licensed source lists (OpenSources.co and Iffy.news) with four computed enrichment signals: domain age (WHOIS/RDAP), web popularity (Tranco Top-1M), fact-check frequency (Google Fact Check Tools API), and threat intelligence (Google Safe Browsing API). The dataset covers 2,672 domains categorized as fake, unreliable, mixed, conspiracy, or satire, each assigned a composite credibility score between 0.0 and 1.0. CRED-1 is designed for on-device deployment in privacy-preserving browser extensions to enable client-side pre-bunking of misinformation at the content delivery stage. The entire pipeline is implemented in Python using only standard library modules and is fully reproducible from publicly available sources. The dataset and pipeline code are released under CC~BY~4.0 and archived on Zenodo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。