研究不同语言下无监督押韵识别所需训练数据量,发现足够数据可让工具超越人类标注一致性。
Training Data Size Sensitivity in Unsupervised Rhyme Recognition
- 用诗歌语料中的重复模式识别押韵,不依赖语言特定特征。
- 在七种语言中,充足数据下模型性能超过人工标注一致性。
- 大模型因缺乏语音表征,在单次学习下表现不佳,适合对语音敏感任务。
押韵看似直观实则复杂:其定义具有历史建构性,学者难以统一分类,人们对于两词是否押韵常有分歧,这给自动化押韵识别与评估带来挑战,尤其在多语言场景下。本文研究使用无需标注的RhymeTagger工具进行无监督押韵识别时,所需训练数据量的影响。该工具基于诗歌语料中的重复模式识别押韵,具有语言无关性。我们在七种语言(捷克语、德语、英语、法语、意大利语、俄语和斯洛文尼亚语)上评估其性能,分析训练规模与语言差异对准确率的影响。为设定合理性能基准,我们对部分手标注诗篇进行人工标注一致性评估,并分析专家标注分歧的主要因素:押韵词间的语音相似度及其在诗中的距离。同时,将RhymeTagger与三种大语言模型在单次学习策略下进行对比。结果表明,当提供足够训练数据时,RhymeTagger的性能持续超越人类标注一致性;而缺乏语音表征的大模型在此任务中表现显著不足。
原文摘要 · Abstract (English)
Rhyme is deceptively intuitive: what is or is not a rhyme is constructed historically, scholars struggle with rhyme classification, and people disagree on whether two words are rhymed or not. This complicates automated rhymed recognition and evaluation, especially in multilingual context. This article investigates how much training data is needed for reliable unsupervised rhyme recognition using RhymeTagger, a language-independent tool that identifies rhymes based on repeating patterns in poetry corpora. We evaluate its performance across seven languages (Czech, German, English, French, Italian, Russian, and Slovene), examining how training size and language differences affect accuracy. To set a realistic performance benchmark, we assess inter-annotator agreement on a manually annotated subset of poems and analyze factors contributing to disagreement in expert annotations: phonetic similarity between rhyming words and their distance from each other in a poem. We also compare RhymeTagger to three large language models using a one-shot learning strategy. Our findings show that, once provided with sufficient training data, RhymeTagger consistently outperforms human agreement, while LLMs lacking phonetic representation significantly struggle with the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。