arXiv:2602.21877cs.CV2026-02中稿 · CVPR

让AI实时指导用户拍出更易记住的照片。

How to Take a Memorable Picture? Empowering Users with Actionable Feedback

  • 用多模态大模型生成自然语言改进建议
  • 在多个基准上显著提升照片记忆度
  • 适合摄影创作者和AI辅助拍摄工具

图像记忆度,即图像被记住的可能性,传统上在计算机视觉中被视为被动预测任务(回归得分)或通过生成方法改变视觉输入以提升记忆可能性。然而这些范式均无法在拍摄时为用户提供支持。本文提出记忆度反馈(MemFeed)新任务,旨在通过自动化模型向用户提供可操作、人类可理解的指导,以增强图像未来回忆概率。我们提出首个专门为此设计的方法 MemCoach,基于多模态大语言模型(MLLMs),无需训练,采用教师-学生引导策略,将模型内部激活对齐至从最少到最易记样本中学习到的记忆模式。为系统评估该任务,我们构建了新基准 MemBench,包含序列对齐的拍照数据与标注的记忆度分数。实验表明,使用多种 MLLM 的 MemCoach 在多个零样本模型上表现更优,证明记忆度不仅可预测,还能被教学与指导,推动研究从单纯预测转向面向创作者的可操作反馈。

原文摘要 · Abstract (English)

Image memorability, i.e., how likely an image is to be remembered, has traditionally been studied in computer vision either as a passive prediction task, with models regressing a scalar score, or with generative methods altering the visual input to boost the image likelihood of being remembered. Yet, none of these paradigms supports users at capture time, when the crucial question is how to improve a photo memorability. We introduce the task of Memorability Feedback (MemFeed), where an automated model should provide actionable, human-interpretable guidance to users with the goal to enhance an image future recall. We also present MemCoach, the first approach designed to provide concrete suggestions in natural language for memorability improvement (e.g., "emphasize facial expression," "bring the subject forward"). Our method, based on Multimodal Large Language Models (MLLMs), is training-free and employs a teacher-student steering strategy, aligning the model internal activations toward more memorable patterns learned from a teacher model progressing along least-to-most memorable samples. To enable systematic evaluation on this novel task, we further introduce MemBench, a new benchmark featuring sequence-aligned photoshoots with annotated memorability scores. Our experiments, considering multiple MLLMs, demonstrate the effectiveness of MemCoach, showing consistently improved performance over several zero-shot models. The results indicate that memorability can not only be predicted but also taught and instructed, shifting the focus from mere prediction to actionable feedback for human creators.

图像记忆度多模态模型摄影建议人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。