发现大模型内部可线性解码作文质量信息,揭示评分机制的可解释性。
From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

- 通过线性探测等方法,发现作文质量信息以线性形式编码在模型表征中。
- 深层网络更擅长处理长作文,且存在与评分强相关的特异性神经元。
- 不同提示和评分标准下仍保持稳定,适合用于提升AI作文评分可解释性。
近期大型语言模型(LLMs)显著推动了自动作文评分(AES)的发展,但其内部评分机制仍不清晰。本文系统分析了八种LLM在两个英文数据集(ASAP++、CSEE)和一个葡萄牙语数据集(ENEM)上的隐藏表征。通过线性探测、跨提示泛化、降维及神经元级分析,发现作文质量信息以线性可解码形式存在于模型表征中,随层深度逐步显现,对提示策略鲁棒,并在不同评分标准间部分迁移。非线性探测仅带来微弱且不一致的提升,表明多数质量信息已线性可读。我们识别出若干“作文评分神经元”,其激活与分数强相关,且对干预敏感。此外,这些神经元的层级分布随作文长度变化,长文更依赖深层表示。研究为大模型中的作文质量结构表征提供了证据,深化了对基于大模型的自动评分系统可解释性的理解。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly understood. In this work, we systematically analyze the hidden representations of eight LLMs across two English essay datasets (ASAP++, CSEE) and one Portuguese dataset (ENEM). Using linear probing, cross-prompt generalization, dimensionality reduction, and neuron-level analyses, we find consistent evidence that essay quality information is encoded in a linearly accessible form within LLM representations. These representations emerge progressively across layers, remain robust across prompting strategies, and partially transfer across essay prompts despite differences in scoring rubrics. In addition, nonlinear probes provide only marginal and inconsistent improvements over linear probes, suggesting that most essay quality information is already linearly decodable. We further identify individual ``essay scoring neurons'' whose activations strongly correlate with essay scores and whose behavior is sensitive to targeted intervention. Moreover, the layer-wise distribution of these neurons systematically shifts with essay length, with longer essays relying more heavily on deeper layers. Overall, our findings provide evidence that LLMs encode structured representations related to essay quality and offer new insights into the interpretability of LLM-based AES systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。