对比开源与闭源大模型在希腊语上的表现,发现各有优劣且有伦理启示。
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
- 比较Llama-70b与GPT-4o mini在7个核心任务的表现,揭示模型差异
- 0-shot作者归属准确率高,暗示预训练数据可能涉及版权问题
- 法律文本三步法优于传统TF-IDF,适合长文本分析场景
低资源语言的自然语言处理长期面临数据匮乏、高资源语言偏见继承及领域适配需求等挑战。本研究以现代希腊语为例,提出三项关键贡献:首先,在七个核心NLP任务上评估开源(Llama-70b)与闭源(GPT-4o mini)大语言模型性能,揭示任务特异性优劣与性能对齐;其次,将作者归属任务重构为评估模型预训练数据使用潜力的工具,0-shot准确率高,提示潜在数据溯源伦理风险;第三,展示法律NLP案例,通过“摘要-翻译-嵌入”(STE)方法对长篇法律文本聚类,效果优于传统TF-IDF。这些成果为推进低资源语言NLP提供了模型评估、任务创新与实际应用的路线图。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) for lesser-resourced languages faces persistent challenges, including limited datasets, inherited biases from high-resource languages, and the need for domain-specific solutions. This study addresses these gaps for Modern Greek through three key contributions. First, we evaluate the performance of open-source (Llama-70b) and closed-source (GPT-4o mini) large language models (LLMs) on seven core NLP tasks with dataset availability, revealing task-specific strengths, weaknesses, and parity in their performance. Second, we expand the scope of Greek NLP by reframing Authorship Attribution as a tool to assess potential data usage by LLMs in pre-training, with high 0-shot accuracy suggesting ethical implications for data provenance. Third, we showcase a legal NLP case study, where a Summarize, Translate, and Embed (STE) methodology outperforms the traditional TF-IDF approach for clustering \emph{long} legal texts. Together, these contributions provide a roadmap to advance NLP in lesser-resourced languages, bridging gaps in model evaluation, task innovation, and real-world impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。