用大模型中间层激活值实现跨提示作文评分,效果优于仅看输出。
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
- 用大模型中间层激活值做评分特征,替代只分析输出。
- 激活值能准确区分不同作文质量,跨提示场景表现稳定。
- 适合想提升自动评分鲁棒性的研究者与教育技术开发者。
自动化作文评分(AES)在跨提示设置中面临评分标准多样性的挑战。以往研究多聚焦于大语言模型(LLMs)的输出以提升评分精度,但我们认为中间层激活同样蕴含重要信息。为验证这一假设,我们评估了LLMs激活在跨提示作文评分任务中的判别能力。具体地,利用激活值训练探测器,并分析不同模型及输入内容对判别力的影响。通过计算不同提示下各特质维度的作文方向,我们考察了大语言模型在不同作文类型和特质上的评价视角变化。结果表明,激活值具备强大的判别能力,且LLMs能根据特质和作文类型调整评价视角,有效应对跨提示场景中评分标准的多样性。
原文摘要 · Abstract (English)
Automated essay scoring (AES) is a challenging task in cross-prompt settings due to the diversity of scoring criteria. While previous studies have focused on the output of large language models (LLMs) to improve scoring accuracy, we believe activations from intermediate layers may also provide valuable information. To explore this possibility, we evaluated the discriminative power of LLMs' activations in cross-prompt essay scoring task. Specifically, we used activations to fit probes and further analyzed the effects of different models and input content of LLMs on this discriminative power. By computing the directions of essays across various trait dimensions under different prompts, we analyzed the variation in evaluation perspectives of large language models concerning essay types and traits. Results show that the activations possess strong discriminative power in evaluating essay quality and that LLMs can adapt their evaluation perspectives to different traits and essay types, effectively handling the diversity of scoring criteria in cross-prompt settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。