arXiv:2506.11243cs.CLcs.AI2025-06被引 3

小模型也能在智能辅导评估中表现优异,证明轻量化模型具备实际应用潜力。

RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?

  • 采用参数少于10亿的轻量级模型,模拟资源受限环境下的研究场景。
  • 在五个赛道中与冠军模型差距最小6.46分,最大仅13.13分,表现接近领先水平。
  • 成果适合计算资源有限的团队或发展中国家机构参考借鉴。

本文介绍我们参与BEA 2025共享任务的RETUYT-INCO方案。出于对全球南方科研机构算力受限的现实考虑,我们主动采用参数少于10亿的小模型。尽管如此,我们的模型在各项任务中仍保持竞争力。根据主办方公布的exact F₁分数,与冠军队伍的差距分别为:第1赛道6.46分、第2赛道10.24分、第3赛道7.85分、第4赛道9.56分、第5赛道13.13分。最小差距为6.46分,最大为13.13分。结果表明,参数小于10亿的模型在这些任务上具备可竞争性,且可在低预算显卡甚至无显卡的设备上运行。

原文摘要 · Abstract (English)

In this paper, we present the RETUYT-INCO participation at the BEA 2025 shared task. Our participation was characterized by the decision of using relatively small models, with fewer than 1B parameters. This self-imposed restriction tries to represent the conditions in which many research labs or institutions are in the Global South, where computational power is not easily accessible due to its prohibitive cost. Even under this restrictive self-imposed setting, our models managed to stay competitive with the rest of teams that participated in the shared task. According to the $exact\ F_1$ scores published by the organizers, the performance gaps between our models and the winners were as follows: $6.46$ in Track 1; $10.24$ in Track 2; $7.85$ in Track 3; $9.56$ in Track 4; and $13.13$ in Track 5. Considering that the minimum difference with a winner team is $6.46$ points -- and the maximum difference is $13.13$ -- according to the $exact\ F_1$ score, we find that models with a size smaller than 1B parameters are competitive for these tasks, all of which can be run on computers with a low-budget GPU or even without a GPU.

轻量模型智能评估算力限制小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。