arXiv:2412.18163cs.CLcs.AI2024-12综述

构建支持印地语与马拉地语的文本匿名化、摘要与拼写检查平台

Survey of Pseudonymization, Abstractive Summarization & Spell Checker for Hindi and Marathi

  • 整合三种NLP功能于统一平台,支持印地语和马拉地语
  • 为使用区域语言的企业与个人用户提供实用工具
  • 填补印度地区语言NLP工具空白,推动本土化应用

印度丰富的语言多样性为技术进步带来了独特挑战与机遇,尤其在自然语言处理(NLP)领域。尽管主流语言的NLP应用已取得显著进展,印度地区语言如马拉地语和印地语仍缺乏充分支持。当前针对印度地区语言的NLP研究尚处起步阶段,具有重要价值。本文旨在构建一个平台,使用户可在英语、印地语和马拉地语中使用文本匿名化、抽象式文本摘要和拼写检查等功能。这些工具主要服务于以印度地区语言为主的企事业单位和个人用户。

原文摘要 · Abstract (English)

India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP applications for widely spoken languages, the regional languages of India, such as Marathi and Hindi, remain underserved. Research in the field of NLP for Indian regional languages is at a formative stage and holds immense significance. The paper aims to build a platform which enables the user to use various features like text anonymization, abstractive text summarization and spell checking in English, Hindi and Marathi language. The aim of these tools is to serve enterprise and consumer clients who predominantly use Indian Regional Languages.

文本匿名化自动摘要拼写检查印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。