arXiv:2502.02167cs.CLcs.IR2025-02被引 1

跨语言新闻网页属性提取,支持六种语言

Multilingual Attribute Extraction from News Web Pages

  • 用MarkupLM和DOM-LM模型在多语言新闻页上微调
  • 六语言数据集达3172页,翻译成英文可提升提取效果
  • 适合做国际新闻数据抓取的研究者与工程师

本文针对多语言新闻网页中自动提取属性的挑战提出解决方案。现有神经网络模型虽在半结构化网页信息抽取中表现优异,但主要应用于电商领域且基于英语预训练,难以直接用于其他语言。为此,我们构建了一个包含3172个标注新闻网页的多语言数据集,涵盖英语、德语、俄语、中文、韩语和阿拉伯语,来自161个网站,已公开于GitHub。我们对最先进的MarkupLM模型进行微调以提取新闻属性,并评估将网页翻译为英语对提取质量的影响。同时,在多语言数据上预训练另一先进模型DOM-LM并微调至该数据集。与现有开源工具相比,两个微调模型均取得更优的抽取指标。

原文摘要 · Abstract (English)

This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high efficacy in extracting information from semi-structured web pages. However, these models are predominantly applied to domains like e-commerce and are pre-trained using English data, complicating their application to web pages in other languages. We prepared a multilingual dataset comprising 3,172 marked-up news web pages across six languages (English, German, Russian, Chinese, Korean, and Arabic) from 161 websites. The dataset is publicly available on GitHub. We fine-tuned the pre-trained state-of-the-art model, MarkupLM, to extract news attributes from these pages and evaluated the impact of translating pages into English on extraction quality. Additionally, we pre-trained another state-of-the-art model, DOM-LM, on multilingual data and fine-tuned it on our dataset. We compared both fine-tuned models to existing open-source news data extraction tools, achieving superior extraction metrics.

多语言信息提取新闻数据跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。