构建古诗专精数据集,用LoRA微调提升翻译与情感理解能力
System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

- 将古诗赏析拆解为释词、释义、情感推断三任务,针对性设计数据集
- 新数据集CCPoetry-49K含49,404对高质量指令-响应,专攻古诗理解
- 基于Qwen2.5-14B用LoRA微调得PoetryQwen,性能提升9.7%至0.757
大语言模型在古典中文翻译和古诗生成领域取得显著进展,但针对古诗精准翻译与情感语义理解的领域研究仍有限。主要挑战在于多数研究将诗歌赏析视为通用任务,忽视其独特性,且高质量领域数据稀缺。为此,本文将任务分解为术语解释、语义解释和情感推理三个子任务,基于多个开源数据集进行清洗与对齐,构建了专用于该领域的《古典诗词指令对数据集》(CCPoetry-49K),包含49,404对高质量指令-响应对。随后,通过低秩适应(LoRA)微调Qwen2.5-14B模型,提出领域专用模型PoetryQwen。在CCL25-Eval Task 5基准测试中,PoetryQwen得分0.757,相较Qwen2.5-14B-Instruct基线(0.690)提升9.7%。结果表明,PoetryQwen显著提升了古诗的精准翻译与情感理解能力。本文提供新数据集与方法论,助力大模型在特定领域的优化。
原文摘要 · Abstract (English)
Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited. The main challenge is that most studies treat the poetic appreciation task as a general-domain problem, neglecting the distinctive features of poetic appreciation, while high-quality and domain-specific datasets are extremely limited. To address this limitation, we decompose the task into three subtasks: term interpretation, semantic interpretation, and emotional inference. Based on multiple open-source datasets, we perform data cleansing and alignment to construct the Classical Chinese Poetry Instruction Pair Dataset (CCPoetry-49K), which comprises 49,404 high-quality instruction-response pairs explicitly optimized for this domain. We then propose a domain-specialized LLM, called PoetryQwen, by applying Low-Rank Adaptation (LoRA) to fine-tune the Qwen2.5-14B model. Experimental results on the CCL25-Eval Task 5 benchmark demonstrate that PoetryQwen achieves a score of 0.757, representing a 9.7% improvement over the Qwen2.5-14B-Instruct baseline (0.690). These findings clearly indicate that PoetryQwen significantly enhances performance in precise translation and emotional understanding of classical poetry. We present new dataset and methodological considerations intended to support the domain-specific optimization of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。