评测诗歌生成图像的多维度基准,揭示模型情感还原能力。
TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation

- 构建1280张诗歌配图数据集,涵盖十维人类标注。
- 发现主流模型难捕捉诗歌隐含情感,图文匹配度不等于意境还原。
- 开源可复用的评分器,支持新模型与新诗体自动评估。
文本到图像(T2I)模型被越来越多地用于表现文学与文化内容,但目前缺乏有效方法衡量图像是否准确传达诗歌意义。该任务具有多重维度:优质插图需视觉合理、忠实于诗歌意象与场景、符合文化与风格特征、无多余文字,并真实传递情感——而最深层的要求,如意象与隐含情绪,并未在文字中明示。现有指标(如CLIPScore、BLIPScore、VQAScore)仅奖励字面文本-图像对应关系,无法判断图像是否成功,更无法解释原因,甚至难以区分最优与最差模型。本文提出TangPoetryBench,一个包含1,280张图像的多维度基准(320首唐代古诗 × 4种先进T2I模型),并获得质量控制的人类标注,覆盖十个评价维度。分析该数据揭示了当前T2I模型的共性与特性优劣,包括对诗歌隐含情感的还原能力。进一步提出PoemAutoEvaluator(PAE),一个开源且基于评分标准的自动评估器,其性能达到强专有裁判(Claude)水平,能泛化至未见过的生成器与第二类诗体(宋词),并使基准无需新增人工标注即可扩展至新图像。本文发布基准、标注与评估工具。
原文摘要 · Abstract (English)
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem. The task is many-sided: a good illustration must be visually sound, faithful to the poem's imagery and scene, culturally and stylistically apt, free of spurious text, and true to its emotion, and its deepest requirements, imagery and especially implicit emotion, are never stated in the words. Existing metrics (CLIPScore, BLIPScore, VQAScore) reward literal text-image correspondence and so cannot tell whether an illustration succeeds, let alone why, or even separate the best model from the worst. We introduce TangPoetryBench, a multi-dimensional benchmark of 1,280 images (320 classical Chinese Tang poems x 4 state-of-the-art T2I models) with quality-controlled human annotations across ten dimensions. Analyzing this data, we reveal the shared and model-specific strengths and weaknesses of current T2I models, including their ability to evoke a poem's implicit emotion. We further introduce PoemAutoEvaluator (PAE), an open, rubric-conditioned evaluator that reaches parity with a strong proprietary judge (Claude), generalizes to an unseen generator and a second poetic tradition (Song Ci), and lets the benchmark scale to new images without fresh human annotation. We release the benchmark, annotations, and evaluator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。