从一张人像图生成三段式电影级构图,提升视觉叙事能力。
ShotCrop$^3$: Cropping Human-Centric Images into Cinematic Triple-Shot Compositions

- 提出三段式构图任务,生成全景、中景、特写三幅画面
- 通过伪标签与强化学习,在1.2千个标注样本上实现定位精度提升2.82倍
- 适合影视海报设计、多场景视觉叙事等创意工作流
以往的美学构图研究通常只生成单一美观裁剪,忽略了从同一场景中构建多个镜头的叙事价值。在实际应用中,多镜头组合对下游创作流程至关重要:商业海报常需不同侧重的多组裁剪(如环境、主体、情感或产品细节),以呈现关键剧情节点。为此,我们提出 extbf{三段式构图(TSC)}任务,从单个人像图像生成一组三镜头——建立镜头、中景和特写——每幅均配有简短镜头描述,支持视觉叙事。为在有限专家标注下学习该任务,我们提出 extbf{ShotCrop},采用三阶段训练:首先进行基于思维链的监督微调以建立基础推理与构图能力;接着利用高置信度伪标签进行半监督微调,进一步增强美学表现力;最后通过面向 extbf{ShotCrop}的组相对策略优化(GRPO-S)与定制化复合奖励进行优化。具体地,我们的伪标签策略融合多模态大模型评分、美学评估与CLIP相似性,保留高质量训练信号。此外,我们构建了包含1200个专家标注测试用例的TSC-Bench基准。值得注意的是,ShotCrop在镜头定位准确率上相较GPT-5平均提升 extbf{2.82}倍。
原文摘要 · Abstract (English)
Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice, multi-shot composition is critical for downstream creative workflows: commercial posters often require multiple crops with different emphases (e.g., context, subject, and emotion/product details) to present key story beats. Therefore, we propose \textbf{Triple-Shot Compositions (TSC)}, a composition task that generates a three-shot set -- establishing, medium, and close-up -- from a single human-centric image, each paired with a brief shot description to support visual narration. To learn TSC with limited expert annotations, we introduce \textbf{ShotCrop} which undergoes a three-stage training process: it first applies Chain-of-Thought supervised fine-tuning to establish basic reasoning and aesthetic shot-cropping skills, then performs semi-supervised fine-tuning with high-confidence pseudo labels to further enhance aesthetic capability, and is finally optimized with Group Relative Policy Optimization for \textbf{ShotCrop} (GRPO-S) using a composite reward tailored for it. Specifically, our pseudo-labeling strategy combines MLLM-based scoring, aesthetic assessment, and CLIP similarity to retain high-confidence training signals. In addition, we present TSC-Bench, a benchmark of 1.2k expert-annotated test cases. Notably, ShotCrop achieves an average improvement of \textbf{2.82} times over GPT-5 in shot localization accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。