arXiv:2608.28695cs.CVcs.IR2026-08中稿 · EMNLP

让AI从零组合图片讲故事,突破传统检索只看单图的局限。

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

论文配图:Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
图 1 · 摘自论文原文
  • 用智能体动态发现图像间关系,组成连贯视觉故事
  • 在10万张图中构建667个验证查询,挑战组合爆炸难题
  • 适合做个性化相册搜索、创意内容生成的研究者

图像检索传统上被当作点对点匹配问题,每个候选图像独立评分。然而,这种原子范式无法捕捉用户在个人照片库中寻求结构化视觉叙事的真实意图。为此,我们提出**图像包组合(IBC)**新范式,将目标从排序单个图像转变为从海量无结构照片池中动态组合连贯的图像包。由于目标包未预定义,IBC面临严重的组合爆炸挑战,需建模不可分解的联合相关性。为此,我们构建了首个IBC基准数据集IBCBench,包含109,467张图像和667个经半自动验证的查询。同时提出**BundleWeaver**,一种基于智能体的框架,将IBC重构为条件式增量超边发现任务。通过大语言模型自适应搜索缺失关系角色,并利用视觉-语言模型进行整包验证,该框架有效导航组合空间。大量实验表明,尽管先进嵌入模型与静态分解重排序范式存在关系盲区,BundleWeaver仍实现显著性能提升,凸显从原子打分转向动态关系组合的必要性。数据集与代码已公开。

原文摘要 · Abstract (English)

Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural relations rather than isolated snapshots. To address this limitation, we introduce **Image Bundle Composition (IBC)**, a novel paradigm that shifts the objective from ranking individual images to dynamically composing cohesive image bundles from a massive, unstructured photo pool. Since target bundles are not predefined, IBC presents a severe combinatorial explosion challenge and demands modeling non-decomposable joint relevance. To establish this paradigm, we construct **IBCBench**, the first IBC benchmark dataset containing 109,467 images and 667 verified queries, built via a semi-automated verification pipeline. Furthermore, we propose **BundleWeaver**, an agentic framework that reformulates IBC as query-conditioned incremental hyperedge discovery. By employing a Large Language Model to adaptively search for missing relational roles and utilizing a Vision-Language Model for whole-bundle verification, BundleWeaver effectively navigates the combinatorial space. Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition. Our dataset and code are available.

图像生成智能体视觉叙事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。