arXiv:2511.21930cs.CL2025-11

对比大模型提示与微调在中文歌词作者识别中的效果,发现类型差异影响显著。

A Comparative Study of LLM Prompting and Fine-Tuning for Cross-genre Authorship Attribution on Chinese Lyrics

  • 构建跨流派中文歌词数据集,对比零样本大模型与微调模型表现
  • 抽象流派(如爱情)识别准确率远低于结构化流派(如民间传统)
  • 建议扩大数据多样性,避免标签不平衡,优化评估方式

我们针对中文歌词作者识别这一缺乏高质量公开数据的领域,提出一项新研究。贡献包括:(1) 构建一个覆盖多流派、平衡的中文歌词新数据集;(2) 开发并微调专用模型,与使用 DeepSeek LLM 的零样本推理进行对比。验证两个核心假设:一是微调模型优于零样本基线;二是性能受流派影响。实验强烈支持第二个假设:结构化流派(如民间传统)的识别准确率显著高于抽象流派(如爱情与浪漫)。第一个假设仅获部分支持:在真实世界数据及复杂流派测试中(Test1),微调提升鲁棒性与泛化能力;但在较小且合成增强的数据集(Test2)中,收益有限或不明确。我们指出 Test2 的设计缺陷(如标签不平衡、词汇差异浅、流派覆盖窄)会掩盖微调的真实效果。本研究建立了首个跨流派中文歌词作者识别基准,强调流派敏感性评估的重要性,并提供公开数据集与分析框架。建议:扩充并多样化测试集,减少对词级数据增强的依赖,平衡各流派作者数量,探索领域自适应预训练以提升性能。

原文摘要 · Abstract (English)

We propose a novel study on authorship attribution for Chinese lyrics, a domain where clean, public datasets are sorely lacking. Our contributions are twofold: (1) we create a new, balanced dataset of Chinese lyrics spanning multiple genres, and (2) we develop and fine-tune a domain-specific model, comparing its performance against zero-shot inference using the DeepSeek LLM. We test two central hypotheses. First, we hypothesize that a fine-tuned model will outperform a zero-shot LLM baseline. Second, we hypothesize that performance is genre-dependent. Our experiments strongly confirm Hypothesis 2: structured genres (e.g. Folklore & Tradition) yield significantly higher attribution accuracy than more abstract genres (e.g. Love & Romance). Hypothesis 1 receives only partial support: fine-tuning improves robustness and generalization in Test1 (real-world data and difficult genres), but offers limited or ambiguous gains in Test2, a smaller, synthetically-augmented set. We show that the design limitations of Test2 (e.g., label imbalance, shallow lexical differences, and narrow genre sampling) can obscure the true effectiveness of fine-tuning. Our work establishes the first benchmark for cross-genre Chinese lyric attribution, highlights the importance of genre-sensitive evaluation, and provides a public dataset and analytical framework for future research. We conclude with recommendations: enlarge and diversify test sets, reduce reliance on token-level data augmentation, balance author representation across genres, and investigate domain-adaptive pretraining as a pathway for improved attribution performance.

作者识别中文歌词大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。