arXiv:2604.14180cs.CLcs.AI2026-04

中文古籍大模型能内知真假却不会说'我不知道',揭示语言模型的隐性认知与显性表达鸿沟。

Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model

  • 用纯古文数据训练318M模型,通过困惑度跳变检测其对历史真伪的内在区分能力。
  • 内部可识别虚构事件(困惑度提升4.24倍),但外部生成中几乎不使用不确定表达(仅3.5%)。
  • 模型不会主动表征无知,需额外训练信号如强化学习才可能学会说'我不确定'。

我们在包含15.6亿词元的纯古典汉语语料上从零开始训练一个318M参数的Transformer语言模型,完全不包含英文字符或阿拉伯数字。通过系统性的分布外测试,我们探究模型是否能区分已知与未知输入,并在生成文本中表达这种区分。结果发现,模型在内部表现出明显的不确定性分离:真实与虚构历史事件间的困惑度比值达2.39倍(p = 8.9e-11,每组n=92),半虚构事件(真实人物+虚构行为)困惑度最高,达4.24倍(p = 1.1e-16),表明其具备超越句法匹配的事实编码能力。然而,在外部生成中,模型从未学会表达不确定性:古典汉语认知标记在分布外问题中的出现率(3.5%)低于分布内(8.3%,p = 0.023),反映训练数据的修辞惯例而非真正的元认知。该现象在三种语言、三种书写系统、八种规模从110M到1.56B的模型中验证:内部事实编码在六种模型中复现,且随规模增长显现;外部无不确定性表达则在所有模型中一致存在。进一步证明,不确定性表达频率完全由训练数据惯例决定——古典汉语模型呈现“谦逊悖论”(已知话题更频繁使用模糊表达),而日语模型几乎从不使用。我们主张,元认知表达(即说“我不知道”)无法仅靠语言建模自然涌现,必须依赖如基于人类反馈的强化学习等显式训练信号。

原文摘要 · Abstract (English)

We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, with zero English characters or Arabic numerals. Through systematic out-of-distribution (OOD) testing, we ask whether the model distinguishes known from unknown inputs, and whether it expresses that distinction in its generated text. We find a clear dissociation between internal and external uncertainty. Internally, the model exhibits a perplexity jump ratio of 2.39x between real and fabricated historical events (p = 8.9e-11, n = 92 per group), with semi-fabricated events (real figures + fictional actions) showing the highest perplexity (4.24x, p = 1.1e-16), demonstrating genuine factual encoding beyond syntactic pattern matching. Externally, however, the model never learns to express uncertainty: classical Chinese epistemic markers appear at lower rates for OOD questions (3.5%) than in-distribution ones (8.3%, p = 0.023), reflecting rhetorical conventions in the training data rather than genuine metacognition. We test both findings across three languages (Classical Chinese, English, Japanese), three writing systems, and eight models from 110M to 1.56B. The internal factual-encoding effect replicates in six of the eight models, emerging with scale (the two smallest Japanese models do not yet separate real from fabricated history), while the external absence of uncertainty expression holds across all eight. We further show that uncertainty expression frequency is determined entirely by training data conventions -- not epistemic states -- with Classical Chinese models showing a "humility paradox" (more hedging for known topics), while Japanese models almost never hedge. We argue that metacognitive expression -- the ability to say "I don't know" -- does not emerge from language modeling alone and requires explicit training signals such as RLHF.

古典汉语元认知语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。