用大模型自动同步多语言表格信息,提升低资源语言数据质量
Leveraging LLM For Synchronizing Information Across Multilingual Tables
- 通过零样本提示分解任务,提升跨语言表格更新的准确性
- 在信息补全任务上效果提升20.58%,信息更新达1.79%准确率提升
- 适合需要跨语言数据对齐的研究者与多语种知识库建设者
当今大量在线信息集中在英语、法语等高资源语言,导致低资源语言内容常滞后或不完整。以维基百科为例,其多语言表格的跨语言同步仍面临挑战。现有基于规则的方法虽有效,但难以应对复杂性和泛化需求。本文探索使用大语言模型(LLM)进行多语言信息同步,采用零样本提示实现可扩展解决方案。我们构建了信息更新数据集(Information Updation),模拟真实场景中更新过时维基百科表格的过程,并评估模型表现。结果表明,单一提示方法效果有限,因此提出任务分解策略以增强一致性与准确性。所提方法优于现有基线,在信息更新任务上提升1.79%,信息补充任务上提升20.58%,展现出在动态更新和丰富跨架构数据方面的强大能力。
原文摘要 · Abstract (English)
The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content in low-resource languages frequently outdated or incomplete. Recent research has sought to improve cross-language synchronization of Wikipedia tables using rule-based methods. These approaches can be effective, but they struggle with complexity and generalization. This paper explores large language models (LLMs) for multilingual information synchronization, using zero-shot prompting as a scalable solution. We introduce the Information Updation dataset, simulating the real-world process of updating outdated Wikipedia tables, and evaluate LLM performance. Our findings reveal that single-prompt approaches often produce suboptimal results, prompting us to introduce a task decomposition strategy that enhances coherence and accuracy. Our proposed method outperforms existing baselines, particularly in Information Updation (1.79%) and Information Addition (20.58%), highlighting the model strength in dynamically updating and enriching data across architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。