对比五款作者提取工具在多语言新闻中的表现,发现读取质量与Trafilatura最稳定。
Author Unknown: Evaluating Performance of Author Extraction Libraries on Global Online News Articles
- 构建跨语言新闻作者手动标注数据集进行评估
- Go-readability和Trafilatura在多语言中表现最一致
- 所有工具在不同语言间结果波动大,需针对性验证
大规模在线新闻内容分析依赖可靠的元数据提取方法。识别网页新闻文章的作者可支持多种研究问题。尽管已有众多现成的作者提取方案,但针对多语言场景的性能比较仍很少。本文构建了一个跨语言新闻作者的手动标注数据集,并用于评估五款现有软件包及一个定制模型的性能。结果显示,Go-readability和Trafilatura在作者提取中最为一致,但所有工具在不同语言间的表现差异显著。这对希望在分析流程中使用作者数据的研究者尤为重要,表明需针对特定语言和地理区域进一步验证结果可靠性。
原文摘要 · Abstract (English)
Analysis of large corpora of online news content requires robust validation of underlying metadata extraction methodologies. Identifying the author of a given web-based news article is one example that enables various types of research questions. While numerous solutions for off-the-shelf author extraction exist, there is little work comparing performance (especially in multilingual settings). In this paper we present a manually coded cross-lingual dataset of authors of online news articles and use it to evaluate the performance of five existing software packages and one customized model. Our evaluation shows evidence for Go-readability and Trafilatura as the most consistent solutions for author extraction, but we find all packages produce highly variable results across languages. These findings are relevant for researchers wishing to utilize author data in their analysis pipelines, primarily indicating that further validation for specific languages and geographies is required to rely on results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。