arXiv:2603.03884cs.CLcs.AI2026-03

构建首个捷克历史文献零样本话题定位基准,评估大模型精准定位话题段落能力。

CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents

  • 基于捷克历史文献构建人工标注话题定位数据集,支持文档与词级评估。
  • 大模型表现差异显著,最强模型接近人类水平,部分模型出现严重定位失败。
  • 适合关注多语言、历史文本理解与零样本推理的研究者使用。

话题定位旨在识别表达给定话题的文本片段,该话题由名称和描述定义。为此,我们引入一个基于捷克历史文献的人工标注基准,包含人工定义的话题及手动标注的文本段落,并支持在文档与词级上的评估。评估基于人类一致性而非单一参考标注。我们对多种大语言模型(LLMs)以及在精简开发数据集上微调的BERT模型进行了评估。结果表明,不同大模型表现差异显著,性能从接近人类的话题检测到明显的段落定位失败不等。尽管规模较小,最优的蒸馏词嵌入模型仍表现出较强竞争力。数据集与评估框架已公开:https://github.com/dcgm/czechtopic。

原文摘要 · Abstract (English)

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined topics together with manually annotated spans and supporting evaluation at both document and word levels. Evaluation is performed relative to human agreement rather than a single reference annotation. We evaluate a diverse range of large language models alongside BERT-based models fine-tuned on a distilled development dataset. Results reveal substantial variability among LLMs, with performance ranging from near-human topic detection to pronounced failures in span localization. While the strongest models approach human agreement, the distilled token embedding models remain competitive despite their smaller scale. The dataset and evaluation framework are publicly available at: https://github.com/dcgm/czechtopic.

话题定位历史文本零样本多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。