arXiv:2510.19585cs.CLcs.AI2025-10Conference of the …

用大模型检测古籍中的拉丁文片段,为历史研究提供新工具

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

  • 构建多模态数据集,标注724页混合语言古籍中的拉丁文
  • 零样本模型可可靠识别拉丁文,但无法真正理解其语义
  • 适合历史语言学与数字人文研究者使用

本文提出一项新任务:从布局多样、语言混杂的历史文献中提取低资源且噪声较大的拉丁文片段。我们构建了一个包含724页标注页面的多模态数据集,并对大型基础模型在该任务上的表现进行了基准测试与评估。结果表明,当前零样本模型已能实现可靠的拉丁文检测,但缺乏对拉丁文的实际语义理解能力。本研究为处理混合语言语料库中的拉丁文建立了全面基线,支持思想史与历史语言学的量化分析。数据集与代码已公开于 https://github.com/COMHIS/EACL26-detect-latin。

原文摘要 · Abstract (English)

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary zero-shot models is achievable, yet these models lack a functional comprehension of Latin. This study establishes a comprehensive baseline for processing Latin within mixed-language corpora, supporting quantitative analysis in intellectual history and historical linguistics. Both the dataset and code are available at https://github.com/COMHIS/EACL26-detect-latin.

历史语言学多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。