arXiv:2603.24413cs.CL2026-03被引 1

让梵文诗歌生成更合韵律且语义连贯。

Pingala: Prosody-Aware Decoding for Sanskrit Poetry Generation

  • 将诗句分组为行,提升语义连贯性10%。
  • 用长词优先策略和SLP1音译,使韵律对齐率提升46%。
  • 提出无参考评估方法,更贴近真实诗歌质量。

梵文诗歌生成需兼顾语义连贯与严格韵律规则。传统将诗句视为整体序列的方法存在局限。本文提出分组行的解码策略——Pingala,通过偏好较长词元,显著提升每行的词形完整性,使语义连贯性提升10%,同时保持韵律一致性。由于梵文采用音素拼写,引入音素感知的SLP1转写方案,使韵律对齐度提高46%,而语义相似度基本不变。针对缺乏标准参考的评估难题,设计基于交叉编码器的无参考评估方法,实现与真实诗歌实例更好的对齐。该方法适用于如Phi-4等指令微调的大语言模型。

原文摘要 · Abstract (English)

Poetry generation in Sanskrit typically requires the verse to be semantically coherent and adhere to strict prosodic rules. In Sanskrit prosody, every line of a verse is typically a fixed length sequence of syllables adhering to prescribed binary patterns of syllable weights. We observe that instead of treating a verse as a monolithic sequence, segmenting them as grouped-lines leads to significant improvement in semantic coherence by 10\% with comparable metrical adherence. Specifically, Pingala, our proposed decoding approach is designed to encourage every line to have well-formed words and our token selection biases the model towards it by preferring longer tokens. Writing in Sanskrit follows phonemic orthography, hence using a phonetically aware transliteration scheme, SLP1, increased the metrical alignment by 46\% with comparable semantic similarity, for a instruction fine-tuned large language models like Phi-4. We also introduce a new approach for reference-free evaluation using cross-encoders which achieved better alignment with true poetry instances.

诗歌生成梵文韵律语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。