教普通人拍出好照片,能说能示范。
PhotoFramer: Multi-modal Image Composition Instruction
- 分三步指导:移动、放大、换视角,自然语言+示例图
- 用真实照片合成训练数据,提升指导准确性
- 适合想学摄影但没经验的普通用户
拍照构图很重要,但多数用户难以拍出好照片。为提供构图指导,我们提出 PhotoFramer,一个支持多模态构图指令的框架。给定一张构图不佳的照片,它先用自然语言说明如何改进,再生成一张构图良好的示例图像。为训练该模型,我们构建了一个大规模数据集。受人类拍照习惯启发,将构图指导拆分为移动、放大和视角变换三类子任务:移动与放大数据来自现有裁剪数据集,视角变换数据通过两阶段流程获得:首先从多视角数据集中采样不同视角的照片对,训练退化模型将优质照片变为劣质照片;其次将该模型应用于专家拍摄的照片,合成劣质图像形成训练对。基于此数据集,我们微调一个能联合处理和生成文本与图像的模型,实现可操作的文本指导搭配可视化示例。大量实验表明,文本指令能有效引导构图,且结合示例优于仅使用示例的基线。PhotoFramer为普及专业摄影经验提供了实用路径。
原文摘要 · Abstract (English)
Composition matters during the photo-taking process, yet many casual users struggle to frame well-composed images. To provide composition guidance, we introduce PhotoFramer, a multi-modal composition instruction framework. Given a poorly composed image, PhotoFramer first describes how to improve the composition in natural language and then generates a well-composed example image. To train such a model, we curate a large-scale dataset. Inspired by how humans take photos, we organize composition guidance into a hierarchy of sub-tasks: shift, zoom-in, and view-change tasks. Shift and zoom-in data are sampled from existing cropping datasets, while view-change data are obtained via a two-stage pipeline. First, we sample pairs with varying viewpoints from multi-view datasets, and train a degradation model to transform well-composed photos into poorly composed ones. Second, we apply this degradation model to expert-taken photos to synthesize poor images to form training pairs. Using this dataset, we finetune a model that jointly processes and generates both text and images, enabling actionable textual guidance with illustrative examples. Extensive experiments demonstrate that textual instructions effectively steer image composition, and coupling them with exemplars yields consistent improvements over exemplar-only baselines. PhotoFramer offers a practical step toward composition assistants that make expert photographic priors accessible to everyday users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。