构建多视角媒体分析套件,提升政治偏见与事实性检测效果
A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis

- 整合新闻源的多维数据:网站图、超链接图、LLM生成图等
- 在2600家媒体上实现最优融合策略,性能超越现有方法
- 适合研究媒体偏见、信息传播与内容可信度的学者使用
新闻媒体以巨大规模影响公众舆论,自动化检测其政治偏见与事实性变得至关重要。然而,该领域仍缺乏统一资源、跨多种方法的全面评估,以及对表示学习和融合策略的系统分析,尤其是在标签稀疏与数据集多样性背景下。此外,鲜有实证研究总结出哪些方法稳定有效、哪些失败及其原因。本文通过四项贡献填补这些空白:第一,提出MBFC-2025,涵盖约2600家媒体的大型标签集;第二,构建ACL-2020(约900家媒体)和MBFC-2025的多视图表示,包括Alexa图、超链接图、LLM生成图、文章文本及维基百科描述;第三,系统评估嵌入视图与融合策略,含基于强化学习的融合变体;第四,大量实验在ACL-2020上达到当前最佳表现,并为MBFC-2025建立强基准。
原文摘要 · Abstract (English)
News outlets shape public opinion at a scale that makes automated detection of political bias and factuality essential. However, the field still lacks unified resources, comprehensive evaluations across diverse approaches, and systematic analyses of the representations and fusion strategies that matter most, especially under label sparsity and dataset diversity. In addition, there is little empirical work reporting broad, observation-driven findings about what consistently works, what fails, and why. We address these gaps through four main contributions. First, we introduce MBFC-2025, a large-scale label set covering approximately 2,600 outlets from Media Bias/Fact Check (MBFC). Second, we construct multiview representations for ACL-2020 (Panayotov et al., 2022), which includes around 900 outlets, as well as for MBFC-2025. These representations span Alexa graphs, hyperlink graphs, LLM-derived graphs, articles, and Wikipedia descriptions. Third, we provide a systematic evaluation and analysis of embedding views and fusion strategies, including a reinforcement learning-based fusion variant. Fourth, we conduct extensive experiments that achieve state-of-the-art results on ACL-2020 and establish strong benchmarks on MBFC-2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。