arXiv:2512.04222cs.CV2025-12被引 3

用大模型当裁判,让图像分解更准更通用。

ReasonX: MLLM-Guided Intrinsic Image Decomposition

  • 用多模态大模型做相对判断,生成强化学习奖励信号
  • 在真实图像上微调,使反照率误差降低9%-25%,深度精度提升46%
  • 不依赖标注数据,适合各种图像分解模型,通用性强

内在图像分解旨在将图像分离为反照率、深度、法线和光照等物理成分。尽管近期基于扩散模型和Transformer的方法得益于合成数据的成对监督,但在多样真实场景中的泛化能力仍受限。我们提出ReasonX,一种新框架,利用多模态大语言模型(MLLM)作为感知裁判,提供相对内在属性比较,并将这些比较作为GRPO奖励,用于在无标签的真实图像上微调内在分解模型。与生成模型的常规强化学习方法不同,该框架通过奖励裁判的相对评估与模型输出的解析关系之间的一致性,来对齐条件内在预测器。ReasonX具有模型无关性,可适配多种内在预测器。在多个基础架构和模态上均取得显著提升,在IIW反照率上实现9%-25%的WHDR降低,在ETH3D上深度精度最高提升46%,展示了MLLM引导的比较监督在连接低层与高层视觉推理方面的潜力。

原文摘要 · Abstract (English)

Intrinsic image decomposition aims to separate images into physical components such as albedo, depth, normals, and illumination. While recent diffusion- and transformer-based models benefit from paired supervision from synthetic datasets, their generalization to diverse, real-world scenarios remains challenging. We propose ReasonX, a novel framework that leverages a multimodal large language model (MLLM) as a perceptual judge providing relative intrinsic comparisons, and uses these comparisons as GRPO rewards for fine-tuning intrinsic decomposition models on unlabeled, in-the-wild images. Unlike RL methods for generative models, our framework aligns conditional intrinsic predictors by rewarding agreement between the judge's relational assessments and analytically derived relations from the model's outputs. ReasonX is model-agnostic and can be applied to different intrinsic predictors. Across multiple base architectures and modalities, ReasonX yields significant improvements, including 9-25% WHDR reduction on IIW albedo and up to 46% depth accuracy gains on ETH3D, highlighting the promise of MLLM-guided comparative supervision to bridge low- and high-level vision reasoning.

图像分解多模态大模型强化学习无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。