arXiv:2606.25763cs.CV2026-06

用大模型实时指导拍照构图和姿势,让手机摄影更智能。

ShutterMuse: Capture-Time Photography Guidance with MLLMs

论文配图:ShutterMuse: Capture-Time Photography Guidance with MLLMs
图 1 · 摘自论文原文
  • 构建双任务评测基准,覆盖拍摄者构图与被摄者姿势建议。
  • 自研模型ShutterMuse在构图决策上表现最优,姿势建议效率高。
  • 数据集含13万张带文字说明的图像,适合开发交互式摄影助手。

真实摄影需要拍摄时的构图与人物姿态指导。现有美学裁剪评测主要针对事后裁剪预测,忽略对被摄对象的指导,导致多模态大语言模型(MLLMs)在拍摄时指导能力的研究不足。为此,我们提出CaptureGuide-Bench评测基准,包含两个互补任务:摄影师侧的构图决策与优化,以及被摄者侧的场景条件化姿态推荐。评估发现:通用型MLLM可进行构图判断但难以精确定位优化区域;专业美学裁剪模型虽能准确定位裁剪区域,但仅限于优化建议,均无法提供可操作的姿态指导。为支持模型发展,我们构建了包含13万样本的CaptureGuide-Dataset,配有文本推理与结构化视觉标注,并开发了统一的ShutterMuse模型,通过监督与强化学习微调。在CaptureGuide-Bench上的实验表明,ShutterMuse在摄影师侧性能最佳,且在被摄者姿态推荐上表现良好,推理成本显著更低,证明了MLLM作为拍摄时交互助手的潜力。

原文摘要 · Abstract (English)

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction and overlook subject-side recommendations, leaving the capture-time guidance capabilities of multimodal large language models (MLLMs) underexplored. To address this gap, we introduce CaptureGuide-Bench, a benchmark with two complementary tasks: photographer-side composition decision and refinement, and subject-side scene-conditioned pose recommendation. Our evaluation reveals limitations: general-purpose MLLMs can make composition decisions but lack precise refinement localization, while specialized aesthetic cropping models localize crops effectively but are limited to refinement; neither provides actionable pose guidance. To support model development, we further construct CaptureGuide-Dataset, comprising 130K samples with textual rationales and structured visual annotations, and develop ShutterMuse, a unified MLLM trained with supervised and reinforcement fine-tuning. Experiments on CaptureGuide-Bench show that ShutterMuse achieves the best overall photographer-side performance among evaluated baselines and competitive subject-side pose recommendation with substantially lower inference cost, demonstrating the potential of MLLMs as interactive assistants for photography during image capture.

摄影辅助多模态大模型实时指导姿态建议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。