构建多语言维基文章与来源数据集,支持跨语言事实核查。
MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- 收录六倍于原版的维基文章与双倍完整网页引用。
- 文章与引用源文本精确对齐,支持定位与溯源分析。
- 专为跨语言、跨时间的事实核查研究设计。
我们提出MegaWika 2,一个大规模多语言维基百科文章及其引用和抓取网络来源的数据集。文章以丰富结构表示,抓取的源文本以精确字符偏移量存储在文章文本中对应引用位置。MegaWika 2相较原始版本扩展了六倍的文章数量和两倍的完整抓取引用。两者均支持报告生成研究;而MegaWika 2特别设计用于支持跨语言、跨时间的事实核查与分析。
原文摘要 · Abstract (English)
We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character offsets of their citations in the article text. MegaWika 2 is a major upgrade from the original MegaWika, spanning six times as many articles and twice as many fully scraped citations. Both MegaWika and MegaWika 2 support report generation research ; whereas MegaWika also focused on supporting question answering and retrieval applications, MegaWika 2 is designed to support fact checking and analyses across time and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。