激活引导个性会影响作文生成与评分,尤其在语文题中影响更大。
Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring
- 用激活向量控制角色性格,实现模型推理时个性化
- 语文开放题受个性影响最大,质量下降达11倍
- 不同性格评分者呈现明显倾向性,适合教育场景研究
基于激活的个性化引导可在推理阶段调整大语言模型行为,但其在教育应用中的影响尚不明确。本研究在ASAP-SAS基准上,针对三种跨越密集型与专家混合架构的语言模型,考察了七种人格特质的激活向量对短答案生成与自动评分的影响。结果显示,个性引导整体降低答案质量,且对开放性英语语言艺术(ELA)题目影响远大于事实性科学题目,解释与论证类任务退化程度最高可达11倍。评分方面,观察到可预测的正负向校准偏移:‘邪恶’和‘无礼’人格评分更严苛,而‘善良’和‘乐观’人格评分更宽松。ELA任务受评分者个性影响是科学任务的2.5-3倍,且专家混合模型的校准偏移比密集模型大约6倍。据我们所知,这是首个系统研究激活引导人格特质在教育生成与评分中作用的工作。结果表明,在教育场景部署个性化模型时,需考虑任务类型与模型架构的校准差异。
原文摘要 · Abstract (English)
Activation-based steering enables inference-time personalization of large language models, but its effects in educational applications are not well understood. We study activation-based persona vectors representing seven character traits in short-answer generation and automated scoring on the ASAP-SAS benchmark, across three language models spanning dense and mixture-of-experts architectures. Persona steering lowers answer quality overall, with much larger effects on open-ended English Language Arts (ELA) prompts than on factual science prompts. Interpretive and argumentative tasks are particularly sensitive, showing up to 11$\times$ larger degradation. On the scoring side, we observe predictable valence-aligned calibration shifts: ``evil'' and ``impolite'' scorers grade more harshly, while ``good'' and ``optimistic'' scorers grade more leniently. ELA tasks are 2.5-3$\times$ more susceptible to scorer personalization than science tasks, and the mixture-of-experts model shows roughly 6$\times$ larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring. Our findings highlight the need for task- and architecture-aware calibration when deploying personalized models in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。