LLMs能捕捉语法、隐喻和语音特征,提升文本体裁识别能力。
LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
- 构建多语言体裁数据集,融合句法树、隐喻数和语音指标。
- 显式提供语言特征可提升分类性能,但效果因任务而异。
- 适合研究语言模型对深层语言结构的理解能力。
大型语言模型在各类语言任务中展现出巨大潜力,但其是否能有效捕捉原始文本中的句法结构、语音线索及韵律模式尚不明确。为探究这一问题,我们基于古腾堡计划(Project Gutenberg)构建了一个新型多语言体裁分类数据集,涵盖英、法、德、意、西、葡六种语言,每项二元任务(诗歌 vs. 小说;戏剧 vs. 诗歌;戏剧 vs. 小说)包含数千条句子。我们为每条数据添加三类显式语言特征:句法树结构、隐喻数量与语音度量指标,以评估其对分类性能的影响。实验表明,尽管基于原始文本或显式特征的LLM分类器均可学习潜在语言结构,但不同特征在各任务中的贡献程度不均,凸显了在模型训练中引入更复杂语言信号的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate remarkable potential across diverse language related tasks, yet whether they capture deeper linguistic properties, such as syntactic structure, phonetic cues, and metrical patterns from raw text remains unclear. To analysis whether LLMs can learn these features effectively and apply them to important nature language related tasks, we introduce a novel multilingual genre classification dataset derived from Project Gutenberg, a large-scale digital library offering free access to thousands of public domain literary works, comprising thousands of sentences per binary task (poetry vs. novel;drama vs. poetry;drama vs. novel) in six languages (English, French, German, Italian, Spanish, and Portuguese). We augment each with three explicit linguistic feature sets (syntactic tree structures, metaphor counts, and phonetic metrics) to evaluate their impact on classification performance. Experiments demonstrate that although LLM classifiers can learn latent linguistic structures either from raw text or from explicitly provided features, different features contribute unevenly across tasks, which underscores the importance of incorporating more complex linguistic signals during model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。