arXiv:2512.21675cs.CV2025-12被引 20

构建统一感知级图像理解框架,提升模型对美感、质量等视觉特征的识别能力。

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

  • 提出分层定义体系与大规模数据集,统一评估图像美感、质量、结构和纹理。
  • 基于领域自适应预训练与任务对齐强化学习,模型在评分与问答任务中表现更优。
  • 可作为文生图生成的奖励模型,适合多模态理解与生成研究者使用。

多模态大语言模型在视觉定位、分割和描述等任务上取得显著进展,但对感知级图像特征的识别能力仍有限。本文提出UniPercept-Bench,一个面向美感、质量、结构和纹理四个关键领域的统一感知级图像理解框架。建立分层定义体系并构建大规模数据集以评估感知级图像理解能力。在此基础上,开发强基线模型UniPercept,采用领域自适应预训练与任务对齐强化学习,实现对视觉评分(VR)与视觉问答(VQA)任务的稳健泛化。UniPercept在感知级图像理解上优于现有多模态大模型,并可作为文生图生成的即插即用奖励模型。本工作定义了多模态大模型时代下的感知级图像理解标准,通过综合性基准与强基线模型,为推进多模态图像感知理解奠定坚实基础。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified framework for perceptual-level image understanding across three key domains: Aesthetics, Quality, Structure and Texture. We establish a hierarchical definition system and construct large-scale datasets to evaluate perceptual-level image understanding. Based on this foundation, we develop a strong baseline UniPercept trained via Domain-Adaptive Pre-Training and Task-Aligned RL, enabling robust generalization across both Visual Rating (VR) and Visual Question Answering (VQA) tasks. UniPercept outperforms existing MLLMs on perceptual-level image understanding and can serve as a plug-and-play reward model for text-to-image generation. This work defines Perceptual-Level Image Understanding in the era of MLLMs and, through the introduction of a comprehensive benchmark together with a strong baseline, provides a solid foundation for advancing perceptual-level multimodal image understanding.

感知理解多模态图像质量文生图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。