arXiv:2505.03638cs.CV2025-05CVPR被引 4

智能引导用户调整手机拍摄角度,自动优化照片构图。

Towards Smart Point-and-Shoot Photography

  • 用CLIP模型构建构图质量评估系统,识别五级构图优劣。
  • 基于32万张带姿态数据的图像训练,实现实时拍摄角度建议。
  • 适合摄影新手,提升手机随手拍的构图水平。

数亿人日常使用智能手机作为点对点拍摄相机,但多数人缺乏构图技巧。传统相机虽能保证对焦和亮度,却无法指导最佳构图。本文提出首个智能点对点拍摄(SPAS)系统,通过实时引导用户调整相机姿态来优化画面。研究构建了包含4000个场景、32万张带相机姿态信息的图像数据集,并开发基于CLIP的构图质量评估(CCQA)模型,采用可学习文本嵌入技术,区分从“差”到“完美”的五级构图质量。进一步设计相机姿态调整模型(CPAM),分步判断当前视角是否可改进,并输出两个姿态调整角度。该模型采用混合专家架构与门控损失函数,实现端到端训练。实验基于公开数据集验证了系统在构图优化上的有效性。

原文摘要 · Abstract (English)

Hundreds of millions of people routinely take photos using their smartphones as point and shoot (PAS) cameras, yet very few would have the photography skills to compose a good shot of a scene. While traditional PAS cameras have built-in functions to ensure a photo is well focused and has the right brightness, they cannot tell the users how to compose the best shot of a scene. In this paper, we present a first of its kind smart point and shoot (SPAS) system to help users to take good photos. Our SPAS proposes to help users to compose a good shot of a scene by automatically guiding the users to adjust the camera pose live on the scene. We first constructed a large dataset containing 320K images with camera pose information from 4000 scenes. We then developed an innovative CLIP-based Composition Quality Assessment (CCQA) model to assign pseudo labels to these images. The CCQA introduces a unique learnable text embedding technique to learn continuous word embeddings capable of discerning subtle visual quality differences in the range covered by five levels of quality description words {bad, poor, fair, good, perfect}. And finally we have developed a camera pose adjustment model (CPAM) which first determines if the current view can be further improved and if so it outputs the adjust suggestion in the form of two camera pose adjustment angles. The two tasks of CPAM make decisions in a sequential manner and each involves different sets of training samples, we have developed a mixture-of-experts model with a gated loss function to train the CPAM in an end-to-end manner. We will present extensive results to demonstrate the performances of our SPAS system using publicly available image composition datasets.

智能拍照构图优化手机摄影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。