arXiv:2609.04755cs.CL2026-09

用深度学习分析古泰米尔诗文注释对,发现模型更懂词序而非内容。

Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs

论文配图:Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs
图 1 · 摘自论文原文
  • 用递归、Transformer等模型学习诗文与注释的向量表示
  • 模型能识别95.5%的词序正确性,但无法复现注释内容
  • 适合语言学、古籍数字化研究者阅读

我们构建了一个包含1,262对诗文-注释(urai)的语料库,来源为五段古典泰米尔文本,涵盖从语法论述到现代改写的不同风格。研究探索表示学习能否恢复其中信息。训练了循环网络、Transformer编码器、孪生匹配网络、mBART式编码器-解码器及仅解码器语言模型,并在相同数据上设置对照实验。TF-IDF作为无训练词汇检索基线表现良好;生成重叠度显示,仅含25个高频注释词的固定字符串优于仅解码器模型。典型相关系数在高斯噪声下达1.000,但令牌-F1仅在0.02-0.20之间波动;编码器-解码器在验证损失上升后仍持续降低训练损失达十六轮。唯一显著结果是:仅解码器模型在112组最小词对中,有107组偏好真实词序(95.5%),但无法再现未见注释内容。我们公开了数据提取与评估流程,原始注释再分发需授权。

原文摘要 · Abstract (English)

We construct a corpus of 1,262 verse-commentary (urai) pairs from five Classical Tamil source sections, ranging from technical grammatical prose to modern paraphrase, and ask what information representation learning can recover. We train recurrent and Transformer encoders, a Siamese-style pair-matching network, an mBART-style encoder-decoder, and a decoder-only language model. Each analysis is interpreted against an appropriate control on the same data. TF-IDF provides a strong no-training lexical retrieval baseline, alongside representation analyses and generation controls for the learned models. A fixed string containing the 25 most frequent commentary words scores higher on generation overlap than the decoder-only model. Canonical correlation reaches 1.000 on Gaussian noise at these sample sizes, token-F1 spans only about 0.02-0.20 on this corpus, and the encoder-decoder continues to lower training loss for sixteen epochs after validation loss has begun to rise. One narrow result remains: the decoder-only model prefers authentic word order in 107 of 112 minimal-pair comparisons (95.5%), but does not reproduce held-out commentary content. We release the extraction and evaluation protocol; redistribution of the source commentaries remains subject to permission.

古籍数字化语言模型泰米尔语表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。