MASCOT提升图文检索中多属性多样性控制,尤其在复合约束下保持高召回率。
MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval

- 将多属性多样性建模为查询感知的资源分配问题,避免传统流形排斥的缺陷。
- 在复合属性约束下,前10名召回率达94.10%,显著优于对比方法的67.63%。
- 适合需要高精度早期结果多样性的实际检索系统,如地理时间双重约束场景。
视觉语言模型在检索语义相关图像方面表现优异,但实际应用中仅靠相关性不足。系统还需在地理、时间等复合属性上实现结果多样化(RD),而精准控制仍具挑战。现有重排序方法如多源确定性点过程(MS-DPP)依赖相似性表示的流形排斥,虽利于广泛探索,但在离散元数据上的多样性降低任务中,早期召回率大幅下降。为此,本文提出MASCOT(模型感知的子模覆盖用于复合属性图文检索)。不同于流形排斥,MASCOT将多属性多样性视为资源分配问题,将属性投影至由查询驱动重要性的软分箱空间。在三个PixelProse多样性降低任务的平均表现中,MASCOT保持前10名召回率88.58%,而MS-DPP仅为67.63%。在复合约束条件下差异更显著:在PP_geo_hour任务中,当时间和地理多样性被同时抑制时,MS-DPP的前10名召回从0.9737降至0.4931,前1名召回跌至0.23;而MASCOT维持前10名召回0.9410,前1名召回0.7202,且多样性指标高于无约束基线。我们不主张全面优越:在综合多样性-相关性得分上,更简单的消融版本在所有三类任务中取得更高调和均值,表明MASCOT的优势集中在复合约束下的非首项召回。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are highly effective in retrieving semantically relevant images. However, in practice, relevance alone is often insufficient. Systems must also achieve Result Diversification (RD) across composite attributes such as geography and time, a task for which precise control remains challenging. Current re-ranking methods, such as Multi-Source Determinantal Point Processes (MS-DPP), address this using manifold-based repulsion over similarity representations. Although this strategy is effective for broad exploration, it exposes a key limitation in manifold-based models: when subjected to diversity-decrease tasks on discrete metadata, they suffer substantial degradation in early-rank recall. To bridge this gap, we introduce MASCOT (Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval). Instead of relying on manifold repulsion, MASCOT formulates multi-attribute diversity as a resource allocation problem, projecting attributes into a soft-binning space weighted by query-driven importance. Averaged across the three PixelProse diversity-decrease tasks, MASCOT preserves an early-rank recall (R@10) of 88.58%, while MS-DPP retains 67.63%. The margin widens under composite constraints: on PP_geo_hour, where temporal and geographic diversity must be suppressed simultaneously, MS-DPP's recall collapses from 0.9737 to 0.4931 and its top-ranked result degrades to R@1 = 0.23, while MASCOT holds R@10 = 0.9410 and R@1 = 0.7202 at a diversity metric above the unconstrained baseline. We do not claim uniform superiority: on aggregate diversity-relevance scores our own simpler ablations attain higher harmonic means on all three decrease tasks, and MASCOT's advantage is specific to recall beyond rank 1 under composite constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。