首个专用于人像美学评估的多模态模型,提升审美判断精准度。
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
- 构建12维美学标准与10.8万张标注人像数据集
- 融合语言建模与回归头,实现整体与细粒度评估
- 引入元投票机制平衡各模块,适合图像审美研究者
人像美学评估(HIAA)长期缺乏系统研究。本文提出首个面向此任务的综合框架,构建首个专用于人像美学评估的数据集HumanBeauty,包含10.8万张高质量人像,其中5万张经严格筛选并由12维美学标准人工标注,其余5.8万张来自公开数据集的系统筛选。基于此数据集,提出HumanAesExpert多模态视觉语言模型,创新设计专家头融合美学子维度知识,并联合使用语言建模(LM)与回归头。引入元投票器(MetaVoter)聚合三头得分,有效平衡各模块能力。大量实验表明,该模型在整体与细粒度评估上均显著优于现有方法。
原文摘要 · Abstract (English)
Image Aesthetic Assessment (IAA) is a long-standing and challenging research task. However, its subset, Human Image Aesthetic Assessment (HIAA), has been scarcely explored. To bridge this research gap, our work pioneers a holistic implementation framework tailored for HIAA. Specifically, we introduce HumanBeauty, the first dataset purpose-built for HIAA, which comprises 108k high-quality human images with manual annotations. To achieve comprehensive and fine-grained HIAA, 50K human images are manually collected through a rigorous curation process and annotated leveraging our trailblazing 12-dimensional aesthetic standard, while the remaining 58K with overall aesthetic labels are systematically filtered from public datasets. Based on the HumanBeauty database, we propose HumanAesExpert, a powerful Vision Language Model for aesthetic evaluation of human images. We innovatively design an Expert head to incorporate human knowledge of aesthetic sub-dimensions while jointly utilizing the Language Modeling (LM) and Regression heads. This approach empowers our model to achieve superior proficiency in both overall and fine-grained HIAA. Furthermore, we introduce a MetaVoter, which aggregates scores from all three heads, to effectively balance the capabilities of each head, thereby realizing improved assessment precision. Extensive experiments demonstrate that our HumanAesExpert models deliver significantly better performance in HIAA than other state-of-the-art models. Project webpage: https://humanaesexpert.github.io/HumanAesExpert/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。