arXiv:2608.26165cs.CLcs.AI2026-08中稿 · AIED 2026

用轻量编码器实现高效精准的创造力自动评估

Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment

论文配图:Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
图 1 · 摘自论文原文
  • 采用小规模BERT+Poly-Encoder结构,降低计算开销
  • 与人工评分相关性达r=0.74,媲美大模型性能
  • 适合教育场景落地,可在普通硬件上运行

自动化创造力评估长期面临资源消耗大或准确率不足的问题。本文提出基于Poly-Encoder的新方法,在包含约18,000条人类评分答题数据的科学创造性思维测试公开数据集上进行微调。该方法利用小型预训练BERT编码器,性能接近微调的大语言模型(LLM),同时显著降低计算需求。实验中,BERT家族模型与不同poly-code数量组合,与人类评分者达到最高r=0.74(95%置信区间[0.73, 0.75]),与资源密集型LLM表现相当。本研究弥合了高性能与计算效率之间的差距,有望在消费级硬件上实现大规模部署。尽管存在一些局限,结果表明Poly-Encoders是实际、可扩展创造力评估的有前景替代方案,尤其适用于教育领域。

原文摘要 · Abstract (English)

Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally efficient and accurate automated creativity assessment. We fine-tuned a Poly-Encoder on a public dataset from the Scientific Creative Thinking Test, comprised of approximately 18,000 human-rated question responses. Our method leverages small pre-trained BERT encoders, achieving performance comparable to fine-tuned Large Language Models while significantly reducing computational demands. Experiments with the BERT-family models and poly-code counts achieved Pearson correlations of up to r = 0.74, 95% CI [0.73, 0.75] with human raters, matching the performance of resource intensive LLMs. This study bridges the gap between high performance and computational efficiency, potentially enabling widespread implementation of automated creativity assessment on accessible consumer-grade hardware. With some limitations, our findings suggest that Poly-Encoders are a promising alternative to LLMs for practical, scalable creativity assessment in various contexts, especially educational.

创造力评估轻量模型BERT自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。