arXiv:2512.12503cs.AI2025-12

为儿童艺术创作设计多维度评价体系,提升AI评估准确性。

KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs

  • 按9个评分维度对1000+幅儿童画作进行专家标注,支持多维评估与反馈。
  • 在Qwen2.5-VL-7B模型上相关性从0.468提升至0.653,尤其改善感知类维度。
  • 适合教育AI、儿童艺术评测及可解释性研究者使用。

多模态大语言模型在视觉-语言任务中表现卓越,但对艺术表达的评估能力仍有限。审美概念本质抽象且开放,且多模态艺术标注数据稀缺。我们提出KidsArtBench,一个包含1000余幅5-15岁儿童作品的基准数据集,由12位教育专家依据9个量规维度标注,并附有专家评语用于形成性反馈。不同于以往仅提供成人图像单标量评分的美学数据集,KidsArtBench聚焦儿童艺术,将多维标注与评论监督结合,支持序数评估与教学反馈。基于此资源,我们提出一种属性特定的多LoRA方法,每个属性对应评分量规中的一个维度(如写实性、想象力),并采用回归感知微调(RAFT)使预测结果对齐序数尺度。在Qwen2.5-VL-7B模型上,相关性从0.468提升至0.653,感知类维度提升最显著,高阶属性差距缩小。结果表明,教师对齐的监督与属性感知训练能生成具有教育意义的评估,为教育AI的持续发展建立严谨测试基准。数据与代码已发布,并附伦理说明。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) show remarkable progress across many visual-language tasks; however, their capacity to evaluate artistic expression remains limited. Aesthetic concepts are inherently abstract and open-ended, and multimodal artwork annotations are scarce. We introduce KidsArtBench, a new benchmark of over 1k children's artworks (ages 5-15) annotated by 12 expert educators across 9 rubric-aligned dimensions, together with expert comments for feedback. Unlike prior aesthetic datasets that provide single scalar scores on adult imagery, KidsArtBench targets children's artwork and pairs multi-dimensional annotations with comment supervision to enable both ordinal assessment and formative feedback. Building on this resource, we propose an attribute-specific multi-LoRA approach, where each attribute corresponds to a distinct evaluation dimension (e.g., Realism, Imagination) in the scoring rubric, with Regression-Aware Fine-Tuning (RAFT) to align predictions with ordinal scales. On Qwen2.5-VL-7B, our method increases correlation from 0.468 to 0.653, with the largest gains on perceptual dimensions and narrowed gaps on higher-order attributes. These results show that educator-aligned supervision and attribute-aware training yield pedagogically meaningful evaluations and establish a rigorous testbed for sustained progress in educational AI. We release data and code with ethics documentation.

儿童艺术多维评估教育AIMLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。