用大模型自动评画作,评分准还出专家级点评。
Fine-Tuning a Large Vision-Language Model for Artwork's Scoring and Critique
- 微调视觉语言模型,一次输出分数和按维度反馈
- 评分相关性超0.97,平均误差仅3.95分
- 适合艺术教育、创意研究中的自动化评估
评估艺术创造力是创造力研究与艺术教育的基础,但传统人工评分(如托伦斯创造性思维测试)在大规模场景下效率低。现有机器学习方法虽具前景,但多依赖图像特征,缺乏解释性反馈。本文提出一种基于多任务学习的框架,通过微调视觉语言模型 Qwen2-VL-7B 实现对人类绘画作品的自动化创造力评估。数据集包含1000幅人类创作画作,每幅在1-100分制上评分,并配有简短的人类描述(内容或作者说明)。两名专家使用五维量规(原创性、色彩、质感、构图、内容)评分并提供书面评论;采用80/20训练测试划分。在视觉编码器输出上添加轻量回归头,使模型能在一次前向传播中预测数值分数并生成符合量规的反馈。通过在系统提示中嵌入结构化量规和作品描述,约束生成文本与量化预测一致。实验显示评分准确率高:皮尔逊相关系数 >0.97,平均绝对误差约3.95。定性评估表明生成反馈与专家评论语义接近(SBERT余弦相似度均值 = 0.798)。该方法连接计算机视觉与艺术评价,为创造力研究与课堂教学反馈提供可扩展工具。
原文摘要 · Abstract (English)
Assessing artistic creativity is foundational to creativity research and arts education, yet manual scoring (e.g., Torrance Tests of Creative Thinking) is labor-intensive at scale. Prior machine-learning approaches show promise for visual creativity scoring, but many rely mainly on image features and provide limited or no explanatory feedback. We propose a framework for automated creativity assessment of human paintings by fine-tuning the vision-language model Qwen2-VL-7B with multi-task learning. Our dataset contains 1000 human-created paintings scored on a 1-100 scale and paired with a short human-written description (content or artist explanation). Two expert raters evaluated each work using a five-dimension rubric (originality, color, texture, composition, content) and provided written critiques; we use an 80/20 train-test split. We add a lightweight regression head on the visual encoder output so the model can predict a numerical score and generate rubric-aligned feedback in a single forward pass. By embedding the structured rubric and the artwork description in the system prompt, we constrain the generated text to match the quantitative prediction. Experiments show strong accuracy, achieving Pearson r > 0.97 and MAE about 3.95 on the 100-point scale. Qualitative evaluation indicates the generated feedback is semantically close to expert critiques (average SBERT cosine similarity = 0.798). The proposed approach bridges computer vision and art assessment and offers a scalable tool for creativity research and classroom feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。