arXiv:2604.13078cs.CL2026-04

构建首个跨语言《罗摩衍那》章节对齐语料库,支持多语言文学比较研究。

IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages

  • 按章节(sarga)对齐印度多语言《罗摩衍那》文本,结构清晰。
  • 已建成英文、马拉雅拉姆文完整语料,印地、泰米尔等六种语言在建。
  • 提供带来源信息的可机器读取格式,适合数字人文与多语言NLP研究。

《罗摩衍那》是南亚和东南亚最具影响力的文学传统之一,两千年来在众多语言和文化中传播。尽管对各地《罗摩衍那》传统的研究广泛,但支持系统性跨语言分析的计算资源仍有限。本文推出IWLV《罗摩衍那》语料库,首次以章节(sarga)为单位,将瓦尔米基版《罗摩衍那》在多种印度语言间对齐。目前语料库包含完整的英文与马拉雅拉姆文层,印地语、泰米尔语、卡纳达语和泰卢固语层正在建设中。数据以结构化JSONL格式发布,并附带明确的来源元数据,适用于比较文学、语料语言学、数字人文及多语言自然语言处理。据我们所知,这是首个具有显式来源信息和机器可读格式的章节对齐多语言《瓦尔米基·罗摩衍那》语料库。

原文摘要 · Abstract (English)

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramayana traditions, computational resources enabling systematic cross-linguistic analysis remain limited. This paper introduces the IWLV Ramayana Corpus, a structured parallel corpus aligning Valmiki's Ramayana across multiple Indian languages at the level of the sarga (chapter). The corpus currently includes complete English and Malayalam layers, with Hindi, Tamil, Kannada, and Telugu layers in active production. The dataset is distributed in structured JSONL format with explicit provenance metadata, enabling applications in comparative literature, corpus linguistics, digital humanities, and multilingual natural language processing. To our knowledge, this is the first sarga-aligned multilingual parallel corpus of the Valmiki Ramayana with explicit provenance metadata and machine-readable format.

文学语料库多语言对齐数字人文梵语研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。