构建首个可控人像构图理解与生成基准,支持精细美学分析与可控生成。
PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

- 基于5万张真实人像,提供多层级构图标注与解释文本。
- 涵盖构图评分、属性推理与问答等任务,支持细粒度理解。
- 适合研究可控图像生成、美学评估与可解释性的人像模型开发者。
人像构图在视觉审美与传播中至关重要,但现有数据集和基准多聚焦粗粒度审美评分或通用图像美学,缺乏对结构化构图分析与明确构图约束下的可控生成研究支持。本文提出PortraitCraft,一个统一的画像构图理解与生成基准。该基准基于约5万张精心筛选的真实人像图像,包含全局构图评分、13种构图属性标注、属性级解释文本、视觉问答对及面向生成的结构化文本描述。在此数据基础上,建立两个互补任务:构图理解(包括评分预测、细粒度属性推理与图像引导问答)与构图感知生成(在显式构图约束下生成图像)。我们定义标准化评估协议,并提供代表性多模态模型的基线结果。PortraitCraft为未来细粒度人像理解、可解释审美评估与可控生成研究提供全面支持。
原文摘要 · Abstract (English)
Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This limits systematic research on structured portrait composition analysis and controllable portrait generation under explicit composition requirements. In this paper, we introduce PortraitCraft, a unified benchmark for portrait composition understanding and generation. PortraitCraft is built on a dataset of approximately 50,000 curated real portrait images with structured multi-level supervision, including global composition scores, annotations over 13 composition attributes, attribute-level explanation texts, visual question answering pairs, and composition-oriented textual descriptions for generation. Based on this dataset, we establish two complementary benchmark tasks for composition understanding and composition-aware generation within a unified framework. The first evaluates portrait composition understanding through score prediction, fine-grained attribute reasoning, and image-grounded visual question answering, while the second evaluates portrait generation from structured composition descriptions under explicit composition constraints. We further define standardized evaluation protocols and provide reference baseline results with representative multimodal models. PortraitCraft provides a comprehensive benchmark for future research on fine-grained portrait understanding, interpretable aesthetic assessment, and controllable portrait generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。