用大模型把英文维基内容适配后迁移到印地语维基,提升内容质量
On the effective transfer of knowledge from English to Hindi Wikipedia
- 利用大模型将英文资料改写为中立风格,再翻译成印地语
- 使印地语维基条目内容量分别提升65%和62%
- 适合想提升低资源语言知识覆盖的项目团队
尽管维基百科是最大的多语言百科全书,但其内容仍存在明显不均衡。高资源语言(如英语)与低资源语言(如印地语)之间存在显著质量差距,许多印地语条目信息不足。为弥合这一差距,我们提出一种轻量级框架,促进英维基到印地语维基的知识有效迁移。若英维基内容过时,框架从外部资源(如英文书籍)提取信息,利用大语言模型的上下文学习能力,将其改写为符合维基百科中立观点(NPOV)风格的内容,再机器翻译为印地语并整合;若英维基内容完整,则直接迁移知识。实验显示,该框架在自动评估和人工评估中分别使印地语维基条目内容量提升65%和62%。
原文摘要 · Abstract (English)
Although Wikipedia is the largest multilingual encyclopedia, it remains inherently incomplete. There is a significant disparity in the quality of content between high-resource languages (HRLs, e.g., English) and low-resource languages (LRLs, e.g., Hindi), with many LRL articles lacking adequate information. To bridge these content gaps, we propose a lightweight framework to enhance knowledge equity between English and Hindi. In case the English Wikipedia page is not up-to-date, our framework extracts relevant information from external resources readily available (such as English books) and adapts it to align with Wikipedia's distinctive style, including its \textit{neutral point of view} (NPOV) policy, using in-context learning capabilities of large language models. The adapted content is then machine-translated into Hindi for integration into the corresponding Wikipedia articles. On the other hand, if the English version is comprehensive and up-to-date, the framework directly transfers knowledge from English to Hindi. Our framework effectively generates new content for Hindi Wikipedia sections, enhancing Hindi Wikipedia articles respectively by 65% and 62% according to automatic and human judgment-based evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。