用网友卡壳提问数据,无监督建模视觉记忆度。
Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- 从Reddit等平台收集卡壳提问,构建无监督记忆信号数据集
- 8.2万视频+描述,支持生成回忆和多模态卡壳检索任务
- 大模型微调后生成记忆描述优于GPT-4o,首次实现跨模态卡壳检索
视觉内容记忆度研究已持续数十年,应用涵盖人类记忆机制理解与内容设计优化。现有研究受限于人工标注记忆度成本高,数据集规模小且仅提供平均记忆得分,缺乏自然开放回忆中的细微记忆信号。本文首次提出大规模无监督数据集,包含超过82,000个视频及对应描述性回忆数据,利用Reddit等平台的tip-of-the-tongue(ToT)检索查询构建。实验表明,该数据集能有效支持回忆生成与ToT检索任务。在该数据集上微调的大视觉语言模型,在生成视觉内容的开放式记忆描述方面超越GPT-4o。此外,通过对比学习策略,我们构建了首个支持多模态ToT检索的模型。本数据集与模型为视觉记忆度研究开辟新路径。
原文摘要 · Abstract (English)
Visual content memorability has intrigued the scientific community for decades, with applications ranging widely, from understanding nuanced aspects of human memory to enhancing content design. A significant challenge in progressing the field lies in the expensive process of collecting memorability annotations from humans. This limits the diversity and scalability of datasets for modeling visual content memorability. Most existing datasets are limited to collecting aggregate memorability scores for visual content, not capturing the nuanced memorability signals present in natural, open-ended recall descriptions. In this work, we introduce the first large-scale unsupervised dataset designed explicitly for modeling visual memorability signals, containing over 82,000 videos, accompanied by descriptive recall data. We leverage tip-of-the-tongue (ToT) retrieval queries from online platforms such as Reddit. We demonstrate that our unsupervised dataset provides rich signals for two memorability-related tasks: recall generation and ToT retrieval. Large vision-language models fine-tuned on our dataset outperform state-of-the-art models such as GPT-4o in generating open-ended memorability descriptions for visual content. We also employ a contrastive training strategy to create the first model capable of performing multimodal ToT retrieval. Our dataset and models present a novel direction, facilitating progress in visual content memorability research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。