arXiv:2512.06662cs.CV2025-12被引 2

通过建模个人看图习惯,让图像描述更符合真人视角。

Personalized Image Descriptions from Attention Sequences

  • 用注意力序列捕捉个体看图顺序和关注点
  • 在4个数据集上平均提升24%的描述质量
  • 适合需要个性化生成的多模态应用

人们观看同一张图像时会关注不同区域、对象和细节,且观察顺序与语言风格各异,导致描述差异显著。现有个性化图像描述模型仅关注语言风格,未利用个体视觉行为。本文提出DEPER(DEscription-PERception persona encoder),通过辅助注意力预测任务学习融合语言风格与观图行为的主体嵌入。轻量适配器将该嵌入与冻结的视觉-语言模型对齐,实现无需重训练的少样本个性化。在涵盖多种观察任务及长短描述的4个数据集中,DEPER平均提升24%,表明建模个性化注意力可生成更贴近人类、质量更高的描述。我们认为理解人的感知方式有助于预测其表达;建模人类感知多样性可提升多模态系统的性能与人类对齐度。

原文摘要 · Abstract (English)

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on linguistic style alone, with no prior work leveraging individual viewing patterns. We address this gap by explicitly modeling personalized viewing behavior as a core factor in description generation. Our method, DEPER (DEscription-PERception persona encoder), learns a subject embedding that captures both linguistic style and viewing behavior, guided by an auxiliary attention-prediction task. A lightweight adapter aligns these embeddings with a frozen vision-language model, enabling few-shot personalization without retraining. Across four datasets spanning diverse viewing tasks and both short and detailed descriptions, DEPER achieves a 24% average improvement, showing that modeling personalized attention produces more human-aligned and high-quality descriptions. We posit that understanding how people see helps predict what they say; modeling human diversity in perception can improve both performance and human alignment in multimodal systems.

图像描述个性化注意力序列多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。