用大模型分析拉丁语作者身份,零样本下表现不错但易受语义干扰。
Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin
- 直接用大模型做拉丁文作者归属,无需复杂特征工程。
- 短文本零样本验证准确率超越传统方法,但结果不稳定。
- 模型决策难解释,适合语言学研究者深入调试使用。
本文评估大型语言模型(LLMs)在教父时期拉丁语文本中的作者归属与验证任务表现。研究表明,即使在无须复杂特征工程的零样本条件下,大模型对短文本仍具备较强的作者验证能力。然而,模型极易受语义误导,其作者分析与决策过程难以调控,这与高资源现代语言研究中报告的结果形成反差。尽管在特定情境下,大模型可超越传统基线,但要获得细致且真正可解释的判断,仍需大量实验调优。
原文摘要 · Abstract (English)
This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-shot authorship verification even on short texts without sophisticated feature engineering. Yet, the models can also be easily "mislead" by semantics. The experiments also demonstrate that steering the model's authorship analysis and decision-making is challenging, unlike what is reported in the studies dealing with high-resource modern languages. Although LLMs prove to be able to beat, under certain circumstances, the traditional baselines, obtaining a nuanced and truly explainable decision requires at best a lot of experimentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。