arXiv:2605.23103cs.CLcs.AI2026-05

用BERT区分明清文集标题中的私人书信与致别序。

A Fine-Tuned BERT Classifier for Personal-Letter Titles in Late-Ming and Early-Qing Collected Works

  • 基于5438个手工标注标题微调中文BERT模型
  • 在33部文集中准确识别约5.5万封书信
  • 已部署至中国人物数据库,助力明清文献研究

本文提出Lepton(书信预测)模型,一种针对古典中文文集目录标题的细调BERT分类器,用于判断标题是私人书信还是易混淆的告别序。该模型在33部明末清初文人的文集中,对5438个手工标注的标题进行微调。目前模型已部署于Hugging Face,并被中国人物数据库(CBDB)用于识别中晚明至清初文集中的约五万五千封书信,支撑了明代书信平台的构建。

原文摘要 · Abstract (English)

I present Lepton (Letter Prediction), a fine-tuned BERT classifier that predicts whether a title in a Classical Chinese wenji table of contents is a personal letter or a closely confusable preface (particularly the farewell-preface). Lepton fine-tunes bert-base-chinese on 5438 hand-labeled wenji titles from thirty-three late-Ming and early-Qing literati. I've deployed the model on Hugging Face and has been used at the China Biographical Database (CBDB) to identify approximately fifty-five thousand letters across mid-Ming through early-Qing wenji, populating the Ming Letter Platform.

古典中文文本分类历史文献BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。