arXiv:2607.12443cs.CLcs.DS2026-07被引 3

用小字母表的计算痕迹实现无需机器模型的语言识别

Language Identification with Succinct Machine-Independent Traces

  • 设计仅依赖语言本身的简洁计算痕迹,不需底层生成机器
  • 使用线性于语言字母表大小的极简符号集完成识别
  • 适合对模型可解释性与低资源学习感兴趣的读者

受大语言模型启发,学界重新关注黄金-安格鲁因语言识别极限模型。近期研究发现,训练字符串的计算痕迹和标注能显著提升学习效果,如带注释代码或思维链标记的文本更易学习。然而现有成果依赖于由自动机理论机器生成的庞大符号痕迹。本文解决两个关键问题:能否使用小字母表的痕迹?能否直接从语言本身定义痕迹而不依赖生成机器?我们证明:对任意语言集合,可构造仅依赖语言字母表大小的线性规模符号集,无需任何额外机器模型即可实现识别极限。该方法在理论上为高效、独立于模型的语言识别提供了新路径。

原文摘要 · Abstract (English)

Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation. Recent papers on this question have proposed looking at computational traces and annotations of training strings as a source of additional power for a learner, reflecting empirical regularities such as the way that commented source code is easier to learn from than arbitrary source code, and text annotated with algorithmically generated chain-of-thought tokens can be easier to learn from than the raw text itself. This recent work has shown positive results for language identification in the presence of such computational traces, but the traces in these positive results come from explicit automata-theoretic machine models that generate the language, where the underlying vocabulary of tokens for the traces is very large. In this paper, we address two fundamental issues left open by this line of work: can we achieve positive results with traces that use only a small alphabet, and can we define traces directly from the language itself, without requiring an underlying machine model that generates it? We establish positive results for both of these questions: for an arbitrary collection of languages, we show how to define computational traces that enable identification in the limit, using an alphabet of tokens that is linear in the size of the alphabet that the languages are defined over, and independent of any other properties of the languages.

语言识别计算痕迹形式语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。