arXiv:2506.18091cs.CL2025-06

对比提示工程与微调模型在捷克语指代消解中的表现

Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution

  • 用提示模板测试大模型,微调小模型专攻指代消解
  • 微调模型最高达88%准确率,远超提示方法的74.5%
  • 适合需要高精度、低资源的捷克语自然语言任务

指代消解在自然语言理解中至关重要,尤其在捷克语这类形态丰富的语言中。本文基于普拉格依存树库构建的数据集,对比了两种现代方法:使用指令微调的大语言模型(如Mistral Large 2、Llama 3)进行提示工程,以及针对捷克语指代消解专门训练的mT5和Mistral模型的微调版本。实验表明,提示方法在少样本场景下表现良好(最高74.5%准确率),但微调模型显著更优,特别是mT5-large达到88%准确率,且所需计算资源更少。我们分析了不同指代类型、先行词距离和语料来源下的性能差异,揭示了两类方法的优劣与权衡。

原文摘要 · Abstract (English)

Anaphora resolution plays a critical role in natural language understanding, especially in morphologically rich languages like Czech. This paper presents a comparative evaluation of two modern approaches to anaphora resolution on Czech text: prompt engineering with large language models (LLMs) and fine-tuning compact generative models. Using a dataset derived from the Prague Dependency Treebank, we evaluate several instruction-tuned LLMs, including Mistral Large 2 and Llama 3, using a series of prompt templates. We compare them against fine-tuned variants of the mT5 and Mistral models that we trained specifically for Czech anaphora resolution. Our experiments demonstrate that while prompting yields promising few-shot results (up to 74.5% accuracy), the fine-tuned models, particularly mT5-large, outperform them significantly, achieving up to 88% accuracy while requiring fewer computational resources. We analyze performance across different anaphora types, antecedent distances, and source corpora, highlighting key strengths and trade-offs of each approach.

指代消解大模型微调捷克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。