提出可同时评估文本对齐与主体保留的低成本图像生成评测方法。
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
- 基于视频推理数据集和图像扰动构建大规模训练集,统一评估双目标。
- 在多个类别上提升6.4分(文本对齐)和5.9分(主体保留)。
- 适合需要高效、可靠评估生成质量的研究者与开发者。
主体驱动的文本到图像生成旨在根据文本描述生成图像,同时保持参考主体图像的视觉特征。尽管该技术在个性化图像生成与视频中角色一致性表达等场景具有广泛应用前景,但其进展受限于缺乏可靠的自动评估手段。现有方法或仅评估单一维度(如文本对齐或主体保留),或与人类判断不一致,或依赖昂贵的API评估。为此,我们提出RefVNLI,一种成本低廉的统一评估指标,在一次运行中同时衡量文本对齐与主体保留能力。该模型基于视频推理基准与图像扰动数据构建的大规模数据集进行训练,在多个基准和主体类别(如*Animal*, *Object*)上表现优于或等同于现有基线,文本对齐最高提升6.4分,主体保留最高提升5.9分。
原文摘要 · Abstract (English)
Subject-driven text-to-image (T2I) generation aims to produce images that align with a given textual description, while preserving the visual identity from a referenced subject image. Despite its broad downstream applicability - ranging from enhanced personalization in image generation to consistent character representation in video rendering - progress in this field is limited by the lack of reliable automatic evaluation. Existing methods either assess only one aspect of the task (i.e., textual alignment or subject preservation), misalign with human judgments, or rely on costly API-based evaluation. To address this gap, we introduce RefVNLI, a cost-effective metric that evaluates both textual alignment and subject preservation in a single run. Trained on a large-scale dataset derived from video-reasoning benchmarks and image perturbations, RefVNLI outperforms or statistically matches existing baselines across multiple benchmarks and subject categories (e.g., \emph{Animal}, \emph{Object}), achieving up to 6.4-point gains in textual alignment and 5.9-point gains in subject preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。