arXiv:2511.21024cs.CV2025-11被引 3

让文字控制照片参数更精准,实现相机参数与语义的统一调节。

CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching

  • 分离摄影师意图与相机参数,通过嵌入编码协同控制
  • 参数变化呈近线性响应,多参数组合自然流畅
  • 适合需要精确调整曝光、白平衡等参数的摄影修图场景

文本引导的扩散模型极大推动了图像编辑与生成。然而,在保持物理一致性的同时实现精确的相机参数控制(如曝光、白平衡、变焦)仍具挑战。现有方法或依赖模糊且纠缠的文本提示,难以精准控制相机参数;或为每项参数训练独立分支,牺牲可扩展性、多参数组合能力及对细微变化的敏感度。为此,我们提出 CameraMaster,一个统一的相机感知图像修图框架。核心思想是显式解耦相机指令,并协同融合两类关键信息流:捕捉摄影师意图的指令表征,以及编码精确相机设置的参数嵌入。CameraMaster 首先使用相机参数嵌入调制指令和内容语义,再通过交叉注意力将调制后的指令注入内容特征,生成强相机敏感的语义上下文。同时,指令与参数嵌入作为条件与门控信号注入时间嵌入,实现去噪过程中的统一、逐层调制,强制语义与参数高度对齐。为训练与评估,我们构建了一个包含 7.8 万张图像-提示对的大规模数据集,标注了相机参数。大量实验表明,CameraMaster 对参数变化呈现单调且近线性响应,支持无缝多参数组合,显著优于现有方法。

原文摘要 · Abstract (English)

Text-guided diffusion models have greatly advanced image editing and generation. However, achieving physically consistent image retouching with precise parameter control (e.g., exposure, white balance, zoom) remains challenging. Existing methods either rely solely on ambiguous and entangled text prompts, which hinders precise camera control, or train separate heads/weights for parameter adjustment, which compromises scalability, multi-parameter composition, and sensitivity to subtle variations. To address these limitations, we propose CameraMaster, a unified camera-aware framework for image retouching. The key idea is to explicitly decouple the camera directive and then coherently integrate two critical information streams: a directive representation that captures the photographer's intent, and a parameter embedding that encodes precise camera settings. CameraMaster first uses the camera parameter embedding to modulate both the camera directive and the content semantics. The modulated directive is then injected into the content features via cross-attention, yielding a strongly camera-sensitive semantic context. In addition, the directive and camera embeddings are injected as conditioning and gating signals into the time embedding, enabling unified, layer-wise modulation throughout the denoising process and enforcing tight semantic-parameter alignment. To train and evaluate CameraMaster, we construct a large-scale dataset of 78K image-prompt pairs annotated with camera parameters. Extensive experiments show that CameraMaster produces monotonic and near-linear responses to parameter variations, supports seamless multi-parameter composition, and significantly outperforms existing methods.

图像修复扩散模型相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。