arXiv:2602.08698cs.CL2026-02

研究印度技术讲座翻译难题,聚焦三大语言的机器翻译挑战。

Challenges in Translating Technical Lectures: Insights from the NPTEL

  • 基于NPTEL平台构建技术讲座语料库,注重专业术语与表达风格保留。
  • 发现不同评估指标对复杂语言特征敏感度差异显著。
  • 适合关注多语言教育科技落地与机器翻译评估的研究者。

本研究探讨了在印度语言(孟加拉语、马拉雅拉姆语、泰卢固语)中应用机器翻译的实践问题与方法论意义,结合新兴翻译工作流与现有评估框架进行分析。选题基于语言多样性考量,呼应印度新教育政策(NEP 2020)对多语言教育技术包容性的要求。研究依托印度最大MOOC平台NPTEL作为语料基础,构建了反映技术概念清晰传达的自然口语语料库,强调保持适当语域与词汇选择的重要性。研究发现,不同评估指标对形态丰富且语义紧凑的语言特征表现出显著的敏感性差异,揭示出表面重叠度度量在真实场景中的局限性。

原文摘要 · Abstract (English)

This study examines the practical applications and methodological implications of Machine Translation in Indian Languages, specifically Bangla, Malayalam, and Telugu, within emerging translation workflows and in relation to existing evaluation frameworks. The choice of languages prioritized in this study is motivated by a triangulation of linguistic diversity, which illustrates the significance of multilingual accommodation of educational technology under NEP 2020. This is further supported by the largest MOOC portal, i.e., NPTEL, which has served as a corpus to facilitate the arguments presented in this paper. The curation of a spontaneous speech corpora that accounts for lucid delivery of technical concepts, considering the retention of suitable register and lexical choices are crucial in a diverse country like India. The findings of this study highlight metric-specific sensitivity and the challenges of morphologically rich and semantically compact features when tested against surface overlapping metrics.

机器翻译多语言教育科技语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。