用大模型生成符合音律的梵文诗,精准度超GPT-4
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
- 用约束解码确保梵文诗符音律规范
- 语法准确率达99.86%,远超GPT-4的31.24%
- 适合对古典诗歌生成感兴趣的学者和开发者
文本生成在大型语言模型推动下取得显著进展,但主要集中在高资源语言。这引发一个根本问题:能否利用这些模型为低资源语言(如梵语)生成结构化诗歌?我们提出Chandomitra,一个将英文输入转化为符合Anushtubh韵律的结构化梵文诗歌的数据集。我们评测了多种开源与闭源模型,并深入研究了约束解码与指令微调等技术。约束解码方法在生成符合音律的梵文诗上达到99.86%的语法准确率,远超GPT-4o(单次提示:31.24%)。最佳指令微调模型在语义连贯性上表现更优,虽语法准确率略低,但人类评估显示其更擅长捕捉诗歌特质。数据与代码已公开。
原文摘要 · Abstract (English)
Text Generation has achieved remarkable performance using large language models. It has also been recently well-studied that these large language models are capable of creative generation tasks but prominently for high-resource languages. This prompts a fundamental question: Is there a way to utilize these (large) language models for structured poetry generation in a low-resource language, such as Sanskrit? We present Chandomitra, an English input to structured Sanskrit Poetry translation dataset, specifically adhering to the Anushtubh meter. We benchmark various open and closed models, and scrutinize specialized techniques such as constrained decoding and instruction fine-tuning, for the proposed task. Our constrained decoding methodology achieves 99.86% syntactic accuracy in generating metrically valid Sanskrit poetry, outperforming GPT-4o (1-shot: 31.24%). Our best-performing instruction-tuned model, on the other hand, performs better in semantic coherence with the English input, at the expense of slightly lower syntactic accuracy. Human evaluation further reveals that instruction fine-tuned model is better able to capture the poetic aspects. Data and Code are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。