用机器学习分析唐诗,找出诗人籍贯的語言痕跡。
Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

- 用詩句字元n-gram和語境特徵預測詩人地域分佈。
- 能準確分辨南北區域,晚期差異最明顯,與歷史背景一致。
- 模型誤判有歷史意義,反映北方文風的權威影響。
我們探究唐代詩人的地理來源是否在作品中留下可檢測的語言痕跡。基於《全唐詩》中每位詩人所有詩作,並通過《中國歷代人物傳記資料庫》(CBDB)將詩人與其行政區(十個唐代道)關聯,構建包含357位詩人的詩人級語料庫,並將籍貫預測設為多分類問題。使用字符n-gram TF-IDF與可解釋的領域特徵(意象、季節、典故),經典與神經模型在粗粒度區域(南/北)預測上達到0.69準確率,遠超0.53的基線,且在細粒度道級別上亦顯著優於隨機猜測。三個發現:(i) 各道間語言距離隨地理距離增加而上升(Mantel r=0.40, p≈0.09,九個道),證實詩歌語言存在距離衰減效應;(ii) 詞彙差異隨時間演變:盛唐時期南北方無差異,晚唐差異最大,符合王朝鼎盛期同質化、後期區域分化之歷史脈絡;(iii) 模型自信錯誤具有歷史意義——初唐時所有誤判均為南方詩人被識別為北方,反映北方官話文風的尊崇地位。進一步顯示,即使使用分層冷凍編碼器的古文版Transformer(GuwenBERT)處理全體文本,其表現僅與簡單TF-IDF相當,二者結合也無提升,表明字符n-gram已充分捕捉區域語言信號。研究結果表明,可解釋的機器學習可作為文學史的假說生成工具。
原文摘要 · Abstract (English)
We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attributed to each author in the Complete Tang Poems (Quan Tang Shi) and linking poets to their administrative circuit of origin via the China Biographical Database (CBDB), we build a poet-level corpus of 357 poets across the ten Tang circuits and frame origin prediction as multi-class classification. Using character $n$-gram TF-IDF together with interpretable domain features (imagery, season, and allusion), classical and neural models predict a poet's broad region (South vs.\ North) at $0.69$ accuracy, well above the $0.53$ majority baseline, and finer circuit-level origin above chance. Beyond classification, three findings emerge. (i) Linguistic distance between circuits grows with geographic distance (Mantel $r=0.40$, $p\approx0.09$ over nine circuits), evidence of a distance-decay effect in poetic language. (ii) The signal interacts with time: South/North separability is at chance in the High Tang and strongest in the Late Tang, consistent with court-driven homogenization at the empire's height followed by regional divergence. (iii) The model's confident errors are historically meaningful -- in the Early Tang, every misclassification is a southern poet read as northern, reflecting the prestige of the northern court idiom. We further show that, when given the whole corpus through a hierarchical frozen-encoder representation, a classical-Chinese transformer (GuwenBERT) only matches -- not beats -- simple TF-IDF, and that combining them adds nothing, indicating that character $n$-grams already capture the regional signal. Our results position interpretable machine learning as a hypothesis generator for literary history.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。