arXiv:2603.11295cs.CL2026-03

用大模型给文本自动断代,发现闭源模型表现更优。

Temporal Text Classification with Large Language Models

  • 测试了多种大模型在时间分类任务中的零样本与少样本表现。
  • 闭源模型在少样本下准确率显著高于开源模型。
  • 微调可提升开源模型性能,但仍不及闭源模型。

语言随时间演变,计算模型可学习这些变化以估计文本发表时间。尽管大型语言模型(LLMs)取得进展,其在自动文本断代(即时间文本分类,TTC)上的表现尚未被系统研究。本研究首次对主流闭源模型(Claude 3.5、GPT-4o、Gemini 1.5)和开源模型(LLaMA 3.2、Gemma 2、Mistral、Nemotron 4)在三个历史语料库(两个英文、一个葡萄牙语)上进行系统评估,涵盖零样本与少样本提示及微调设置。结果表明,闭源模型表现优异,尤其在少样本条件下;微调可显著提升开源模型性能,但整体仍无法达到闭源模型水平。

原文摘要 · Abstract (English)

Languages change over time. Computational models can be trained to recognize such changes enabling them to estimate the publication date of texts. Despite recent advancements in Large Language Models (LLMs), their performance on automatic dating of texts, also known as Temporal Text Classification (TTC), has not been explored. This study provides the first systematic evaluation of leading proprietary (Claude 3.5, GPT-4o, Gemini 1.5) and open-source (LLaMA 3.2, Gemma 2, Mistral, Nemotron 4) LLMs on TTC using three historical corpora, two in English and one in Portuguese. We test zero-shot and few-shot prompting, and fine-tuning settings. Our results indicate that proprietary models perform well, especially with few-shot prompting. They also indicate that fine-tuning substantially improves open-source models but that they still fail to match the performance delivered by proprietary LLMs.

文本断代大模型时序分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。