arXiv:2510.01233cs.CL2025-10

用算法识别泰卢固语诗歌格律,助力濒危文化传承

Computational Social Linguistics for Telugu Cultural Preservation: Novel Algorithms for Chandassu Metrical Pattern Recognition

  • 构建首个数字框架,融合社区知识与计算方法分析诗歌韵律
  • 91.73%准确率识别格律模式,符合传统文学评价标准
  • 适合数字人文、文化保护与计算社会科学研究者参考

本研究提出一种计算社会科学方法,用于保护泰卢固语昌达苏(Chandassu)——这一承载数百年集体文化智慧的格律诗歌传统。我们建立了首个综合数字框架,用于分析泰卢固语韵律模式,连接传统社区知识与现代计算技术。该框架通过4,651条标注诗歌构建协作数据集,经专家验证的语言模式,以及文化导向的算法设计实现。包含三部分:AksharamTokenizer(韵律感知分词)、LaghuvuGuruvu Generator(轻重音节分类)和PadyaBhedam Checker(自动模式识别)。算法在提出的Chandassu Score上达到91.73%准确率,评估指标符合传统文学标准。研究表明,计算社会科学可有效保存濒危文化知识体系,并促进文学遗产相关的新型集体智能发展。该方法为以社区为中心的文化保护提供范式,支持数字人文与社会感知计算系统的广泛应用。

原文摘要 · Abstract (English)

This research presents a computational social science approach to preserving Telugu Chandassu, the metrical poetry tradition representing centuries of collective cultural intelligence. We develop the first comprehensive digital framework for analyzing Telugu prosodic patterns, bridging traditional community knowledge with modern computational methods. Our social computing approach involves collaborative dataset creation of 4,651 annotated padyams, expert-validated linguistic patterns, and culturally-informed algorithmic design. The framework includes AksharamTokenizer for prosody-aware tokenization, LaghuvuGuruvu Generator for classifying light and heavy syllables, and PadyaBhedam Checker for automated pattern recognition. Our algorithm achieves 91.73% accuracy on the proposed Chandassu Score, with evaluation metrics reflecting traditional literary standards. This work demonstrates how computational social science can preserve endangered cultural knowledge systems while enabling new forms of collective intelligence around literary heritage. The methodology offers insights for community-centered approaches to cultural preservation, supporting broader initiatives in digital humanities and socially-aware computing systems.

文化保护诗歌格律计算社会学数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。