arXiv:2605.11242cs.CLcs.AI2026-05被引 1

用元提示技术自动适配德国作文评分标准,提升模型泛化能力。

RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German

  • 基于训练数据生成定制化评分提示,动态适应不同评分标准。
  • 在未见答案的三分类任务中获0.729的QWK,排名第六。
  • 适合需要快速适配新评分体系的自动化写作评估场景。

本文介绍我们团队在BEA 2026共享任务“基于评分量规的德语简答评分”中的参与情况。我们参加了第1赛道(未见答案三分类)、第3赛道(未见答案二分类)和第4赛道(未见问题二分类)。由于这些赛道需根据特定评分量规对学生的简短回答进行评分,我们针对任务的动态性提出了名为元提示(Meta-prompting)的方法:利用大语言模型(LLM)从训练集中提取样例,生成自定义评分提示,再用于评判新答案。此外,我们还结合了传统机器学习、开源LLM微调及多种提示技巧。官方结果表明,在第1赛道中,我们以0.729的QWK位列8支队伍中的第6;第3赛道中,以0.674的QWK位居9支队伍中的第4;第4赛道中,以0.49的QWK排名8支队伍中的第4。

原文摘要 · Abstract (English)

In this paper, we present the RETUYT-INCO participation at the BEA 2026 shared task "Rubric-based Short Answer Scoring for German". Our team participated in track 1 (Unseen answers three-way), track 3 (Unseen answers two-way) and track 4 (Unseen questions two-way). Since these tracks required scoring short student answers using specific rubrics, we looked for ways to handle the changing nature of the task. We created a method called Meta-prompting. In this approach, an LLM creates a custom prompt based on examples from the Train set. This prompt is then used to grade new student answers. Along with this method, we also describe other approaches we used, such as classic machine learning, fine-tuning open-source LLMs, and different prompting techniques. According to the official results, our team placed 6th out of 8 participants in Track 1 with a QWK of 0.729. In Track 3, we secured 4th place out of 9 with a QWK of 0.674, and we also placed 4th out of 8 in Track 4 with a QWK of 0.49.

评分系统元提示大模型应用德语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。