arXiv:2601.13105cs.CL2026-01被引 1

用微调+检索增强,让大模型更准识别英语双宾语结构。

Leveraging Lora Fine-Tuning and Knowledge Bases for Construction Identification

  • 用LoRA微调Qwen3-8B,结合检索增强生成框架。
  • 在英国国家语料库上分类准确率显著优于原模型和纯理论系统。
  • 模型从机械匹配转向语义理解,错误分析揭示认知转变。

本研究通过将基于LoRA的大型语言模型微调与检索增强生成(RAG)框架结合,探索自动识别英语双宾语结构的方法。在英国国家语料库的标注数据上进行了二分类任务。结果表明,经LoRA微调的Qwen3-8B模型显著优于原始的Qwen3-MAX模型以及仅依赖理论的RAG系统。详细错误分析显示,微调使模型的判断从表层形式匹配转向更基于语义的理解。

原文摘要 · Abstract (English)

This study investigates the automatic identification of the English ditransitive construction by integrating LoRA-based fine-tuning of a large language model with a Retrieval-Augmented Generation (RAG) framework.A binary classification task was conducted on annotated data from the British National Corpus. Results demonstrate that a LoRA-fine-tuned Qwen3-8B model significantly outperformed both a native Qwen3-MAX model and a theory-only RAG system. Detailed error analysis reveals that fine-tuning shifts the model's judgment from a surface-form pattern matching towards a more semantically grounded understanding based.

自然语言处理模型微调知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。