首个面向现代梵语文本的摘要生成研究,揭示低资源语言挑战
Abstractive Text Summarization for Contemporary Sanskrit Prose: Issues and Challenges
- 构建梵语抽象摘要流程,解决数据与模型训练难题
- 系统梳理从数据收集到推理的全流程挑战
- 为低资源屈折语研究提供可复用范式
本论文针对现代梵语文本开展抽象式文本摘要研究。第一章介绍研究动机、核心问题与理论框架。梵语属低资源屈折语言,本文核心研究问题是:如何克服梵语抽象摘要开发中的挑战?围绕四个主题提出子问题。第二章综述相关研究。第三章聚焦数据准备,解决语言模型与摘要模型训练中的数据采集与预处理难题。第四章报告模型训练与推理结果。研究首次建立梵语抽象摘要流程,系统呈现各阶段面临的具体挑战,并对各主题下的子问题作出回应,最终解答核心问题。
原文摘要 · Abstract (English)
This thesis presents Abstractive Text Summarization models for contemporary Sanskrit prose. The first chapter, titled Introduction, presents the motivation behind this work, the research questions, and the conceptual framework. Sanskrit is a low-resource inflectional language. The key research question that this thesis investigates is what the challenges in developing an abstractive TS for Sanskrit. To answer the key research questions, sub-questions based on four different themes have been posed in this work. The second chapter, Literature Review, surveys the previous works done. The third chapter, data preparation, answers the remaining three questions from the third theme. It reports the data collection and preprocessing challenges for both language model and summarization model trainings. The fourth chapter reports the training and inference of models and the results obtained therein. This research has initiated a pipeline for Sanskrit abstractive text summarization and has reported the challenges faced at every stage of the development. The research questions based on every theme have been answered to answer the key research question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。