大模型无需专门训练即可高效实现印度语言转写。
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
- 用通用大模型直接转写印度语言,省去专门训练
- GPT-4o在十种语言上平均准确率超专用模型
- 适合多语言场景下快速部署的开发者
转写是将文本从一种文字转换为另一种的重要任务,在语言多样性的印度尤为关键。尽管已有专用模型如IndicXlit表现优异,但近期大语言模型的发展表明,通用模型可能无需任务特训即可胜任此任务。本研究系统评估了GPT-4o、GPT-4.5、GPT-4.1、Gemma-3-27B-it和Mistral-Large等主流大模型,对比IndicXlit在十种主要印度语言上的表现。实验基于Dakshina和Aksharantar标准数据集,采用Top-1 Accuracy与字符错误率(CER)进行评估。结果表明,GPT系列模型普遍优于其他大模型及IndicXlit;对GPT-4o进行微调后,在特定语言上性能显著提升。深入的错误分析与噪声环境下的鲁棒性测试进一步揭示了大模型相较于专用模型的优势,证明基础模型在极少额外开销下可有效支持多种专业应用。
原文摘要 · Abstract (English)
Transliteration, the process of mapping text from one script to another, plays a crucial role in multilingual natural language processing, especially within linguistically diverse contexts such as India. Despite significant advancements through specialized models like IndicXlit, recent developments in large language models suggest a potential for general-purpose models to excel at this task without explicit task-specific training. The current work systematically evaluates the performance of prominent LLMs, including GPT-4o, GPT-4.5, GPT-4.1, Gemma-3-27B-it, and Mistral-Large against IndicXlit, a state-of-the-art transliteration model, across ten major Indian languages. Experiments utilized standard benchmarks, including Dakshina and Aksharantar datasets, with performance assessed via Top-1 Accuracy and Character Error Rate. Our findings reveal that while GPT family models generally outperform other LLMs and IndicXlit for most instances. Additionally, fine-tuning GPT-4o improves performance on specific languages notably. An extensive error analysis and robustness testing under noisy conditions further elucidate strengths of LLMs compared to specialized models, highlighting the efficacy of foundational models for a wide spectrum of specialized applications with minimal overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。