无需标注数据,用无监督方法让古语文本生成高质量句子向量。
From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

- 用TSDAE和对比学习两种无监督方法,将专用词级模型转为语料特定的句向量编码器。
- 在2935个专家验证的圣经重用案例上,检测与对应检索均超越所有基线模型。
- 仅需4-8千条原始文本,几秒训练即达最优,适合处理噪音文本与跨作者分析。
自动文本识别(ATR)为数字人文提供了大量未结构化、异质且常含噪声的古语文本。下游任务如文本重用识别、对齐与语义搜索依赖句子嵌入,但现有方法在古语文本上表现不佳:通用多语言编码器性能差,专用语言模型产生非均匀表示空间,且缺乏标注的相似性数据。本文研究两种完全无监督策略——基于重建的TSDAE与对比句向量学习(CSE),利用原始句子将专用词级模型转化为语料特定的句编码器。以教父文献中圣经重用(拉丁语与古希腊语,共2,935个专家验证平行文本,来自奥古斯丁、杰罗姆与亚他那修)为案例,将重用识别分解为二分类检测与对应关系检索两个独立任务,对比多种基线(多语言、专用、蒸馏、有监督微调模型)以及模拟HTR误差与抄写缩写的噪声数据。结果表明,适配后的编码器在两项任务上均优于所有基线,且表现互补:TSDAE在大规模领域内语料下更优,而CSE在检索任务中领先,仅需4-8k条原始领域内句子即可达到最佳性能(在笔记本电脑GPU上训练数秒),并能跨作品、跨作者迁移,即使在重新直接训练于降噪后文本时仍有效。通过UMAP图谱揭示了两种策略的几何特性与其性能提升的关系,完整流程——分段、微调、跨语料语义搜索——已通过在线工具Paraphrasis开放给非专家用户。
原文摘要 · Abstract (English)
Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reuse in patristic literature (2,935 expert-verified parallels in Latin and Ancient Greek, from Augustine, Jerome, and Athanasius), we decompose reuse identification into two separately evaluated tasks-binary detection and correspondence retrieval-and benchmark the adapted encoders against multilingual, specialized, distilled, and supervised fine-tuned baselines, as well as on artificially noised data simulating HTR artifacts and scribal abbreviations. The adapted encoders outperform all baselines on both tasks, with complementary profiles: TSDAE leads detection given a large in-domain corpus, while CSE leads retrieval, reaches its optimum with as few as 4-8k raw in-domain sentences-a few tens of seconds of training on a laptop GPU-and transfers across works and authors, including to noisy post-ATR text when retrained directly on it. UMAP atlases relate the geometric effect of each strategy to the measured gains, and the full pipeline-segmentation, fine-tuning, cross-corpus semantic search-is made available to non-specialists through the online tool Paraphrasis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。