arXiv:2412.04497cs.CLcs.AI2024-12被引 60

用大模型破解冷门语言研究难题,助力文化传承

Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research

  • 利用大模型分析低资源语言的语法与文本特征
  • 揭示大模型在历史文献与文学研究中的应用潜力
  • 强调跨学科合作与定制化模型的重要性

低资源语言承载着人类历史、文化演进与思想多样性,但因数据稀缺与技术限制,其研究与保护面临严峻挑战。近年来的大语言模型(LLMs)为突破这些瓶颈提供了变革性机遇,推动语言学、历史学与文化研究方法创新。本文系统评估了LLMs在低资源语言研究中的应用,涵盖语言变异、历史文献、文化表达与文学分析。通过分析技术框架、现有方法及伦理问题,识别出数据可及性、模型适应性与文化敏感性等关键挑战。鉴于低资源语言蕴含丰富的文化与历史价值,本文倡导跨学科协作与定制化模型开发,以促进该领域研究发展。研究强调将人工智能融入人文学科,有助于保存和理解人类的语言与文化遗产,推动全球智力多样性保护。

原文摘要 · Abstract (English)

Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face critical challenges, including data scarcity and technological limitations, which hinder their comprehensive study and preservation. Recent advancements in large language models (LLMs) offer transformative opportunities for addressing these challenges, enabling innovative methodologies in linguistic, historical, and cultural research. This study systematically evaluates the applications of LLMs in low-resource language research, encompassing linguistic variation, historical documentation, cultural expressions, and literary analysis. By analyzing technical frameworks, current methodologies, and ethical considerations, this paper identifies key challenges such as data accessibility, model adaptability, and cultural sensitivity. Given the cultural, historical, and linguistic richness inherent in low-resource languages, this work emphasizes interdisciplinary collaboration and the development of customized models as promising avenues for advancing research in this domain. By underscoring the potential of integrating artificial intelligence with the humanities to preserve and study humanity's linguistic and cultural heritage, this study fosters global efforts towards safeguarding intellectual diversity.

大模型低资源语言人文学科跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。