arXiv:2607.11873cs.CLcs.LG2026-07

检验教学反馈分类协议在跨语言和模型演进下的稳定性与迁移能力。

A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

  • 用三类表示方法重测原西班牙语数据,验证协议鲁棒性
  • 新模型在主题分类上表现更优,但情感判断无显著优势
  • 跨语言迁移至英文仍有效,适合教育评估系统部署

高校收集的开放性教学评价反馈远超可读数量。此前研究提出一种经验证的分类协议,基于标注指南、标注者一致性测量、分层交叉验证及西班牙语机构语料库的预留测试,采用冻结编码器设计。当前两个问题限制其复用:固定于2019年冻结嵌入的协议是否仍具竞争力,以及能否跨语言迁移。本研究在原西班牙语数据上,以三种表示生成方式重新运行该协议:稀疏词法特征、冻结Transformer嵌入、提示式大语言模型;并将情感任务迁移至英语,使用45,000条平衡评论语料,对照带方面标注的教育数据集进行验证。将成对比较视为描述性分析,结果表明协议具有持久性:2026年前沿模型在最困难的西班牙语主题分类任务中达到最高F1,但在情感分类上未优于低成本模型,且在英语上与低成本模型无描述性差异,因此模型选择是部署决策,而非方法本身的属性。

原文摘要 · Abstract (English)

Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.

教学评估跨语言迁移文本分类协议验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。