arXiv:2607.16330cs.CV2026-07中稿 · CVPR

用多模态大模型评估书法笔画质量并生成教学反馈

Local Brushstroke Quality Assessment via Vision-Language Feedback

论文配图:Local Brushstroke Quality Assessment via Vision-Language Feedback
图 1 · 摘自论文原文
  • 用多模态大模型对比书法前后图像,打分评估笔画质量
  • GPT-4o得分最准(平均绝对误差0.885),但与专家评分相关性不显著
  • 引入检索增强生成反而降低准确率,说明文本规则注入有副作用

本文研究多模态大模型能否评估书法中局部笔画质量,并生成具有教育价值的自然语言反馈。我们构建了一个评估框架,使用三种多模态大模型(GPT-4o、Claude Sonnet 4、Gemini 2.5 Flash)对书法作品的前后对比图像进行五级序数评分,并与三位书法专家的评分进行比较。此外,还初步考察了Claude的检索增强生成(RAG)变体。结果表明,所有模型均达到可接受的绝对评分精度(平均绝对误差,MAE),其中GPT-4o表现最佳(MAE = 0.885)。然而,各模型与人类专家评分之间的整体秩相关性均未达统计显著水平(Kendall's tau)。对生成理由的词汇分析揭示了各模型特有的评价偏见。值得注意的是,RAG虽提升了秩相关性,却恶化了绝对准确性,这一负向结果提示基于文本规则注入存在潜在风险。

原文摘要 · Abstract (English)

This paper investigates whether multimodal LLMs can evaluate local brushstroke quality in calligraphy and generate educationally useful natural language feedback. We construct an evaluation framework in which three multimodal LLMs (GPT-4o, Claude Sonnet 4, and Gemini 2.5 Flash) assess before-after image pairs of calligraphic works using a five-point ordinal scale, and compare their outputs against scores assigned by three expert calligraphers. We additionally examine a Retrieval-Augmented Generation (RAG) variant of Claude as a preliminary condition. Results show that all models achieve useful levels of absolute score accuracy (MAE), with GPT-4o performing best (MAE = 0.885). However, none of the models produce statistically significant overall rank correlations with human experts (Kendall's tau). Vocabulary analysis of generated rationales reveals characteristic evaluative biases in each model, and RAG is shown to improve rank correlation while worsening absolute accuracy, constituting an important negative result for text-based rule injection.

多模态大模型书法评估生成反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。