打造高保真梵文颂诗语音系统,精准还原韵律与发音细节。
Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit
- 基于现有流匹配模型,加入梵语正字与音系处理模块。
- 通过韵律识别与参考片段匹配,实现接近4.6分专家评分的语音质量。
- 适合梵文研究、宗教诵读及语音合成领域开发者使用。
我们提出Vagdhenu,一个关注韵律(vrutta)的梵文诗句到吟诵的文本转语音系统:将格律诗准确转化为高保真吟诵。本文为经验报告,非新架构。采用现成的流匹配TTS骨干网络和大规模神经声码器,并添加梵文吟诵所需组件:前端通过卡纳达拼写避免德瓦纳加里拼音导致的印地语式元音删除;另一前端遵循细微梵语音系规则(如visarga sandhi的jihvamuliya与upadhmaniya变体、alpaprana与mahaprana的送气对比、齿音、卷舌音与腭音擦音的区分);以及韵律感知机制,可检测诗体并按半参考规则匹配精确参照片段。报告关键负结果:在自填充流匹配骨架中,文本侧韵律调节器在结构上无效,因模型从上下文频谱恢复音高,嵌入无法获得梯度;仅参考片段与声音引导重训是有效的韵律控制手段。对比四个家族(StyleTTS2、VITS2、Matcha-TTS与流匹配骨架),每个早期家族在连音或韵律上达到瓶颈,而五小时克隆样本突破上限,专家主观评分接近4.6。系统已部署两个版本:32章共5183句的视频数据集(约17.5小时)和覆盖12部著作共约18000句的音频应用。我们发布前端代码、推理与训练代码、权重、单说话人吟诵数据集及交互式演示。
原文摘要 · Abstract (English)
We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new architecture. We take an off-the-shelf flow-matching TTS backbone and a large-scale neural vocoder, and add the components a faithful Sanskrit chant pipeline needs: a frontend that routes Sanskrit through Kannada orthography to avoid the Hindi-style schwa deletion that Devanagari triggers in Indic models; a frontend that obeys subtle Sanskrit phonology (visarga sandhi with its jihvamuliya and upadhmaniya allophones, the aspiration contrast of alpaprana and mahaprana, and the dental, retroflex, and palatal sibilants kept distinct); and a vrutta-aware mechanism that detects the meter and picks an exactly matched reference under a half-reference rule. We report a negative result that shaped the system: in a self-infilling flow-matching backbone, a text-side prosody conditioner is architecturally inert, because the model recovers pitch from the context mel and the embedding gets no gradient; the reference clip and a voice-steering retrain are the only working prosody levers. We also report a comparative lineage across four families (StyleTTS2, VITS2, Matcha-TTS, and the flow-matching backbone), where each earlier family hit a ceiling on conjuncts or prosody that a five-hour clone cleared at an expert MOS near 4.6. The system shipped two deployments: a 32-chapter, 5183-verse video corpus (about 17.5 hours) and an audio app covering about 18000 verses across 12 books. We release the frontend, inference and training code, weights, a single-speaker chant dataset, and an interactive demo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。