让图文检索结果按上下文自动优化多属性多样性。
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
- 用多源确定性点过程建模,统一处理多种属性的多样性。
- 在多个数据集上显著提升复合属性的多样性表现。
- 适合需要精准控制图像多样性的实际应用。
结果多样化(RD)是提升图文检索实用效率的关键技术。传统方法仅关注图像外观的多样性,但多样性指标及其理想值随应用场景而异,限制了其应用范围。本文提出新任务CDR-CA(复合属性的上下文多样性精炼),旨在根据应用上下文对多种属性的多样性进行优化。为此,我们提出多源确定性点过程(MS-DPP),将标准确定性点过程(DPP)扩展至多源场景。通过流形表示构建统一相似性矩阵,建模为单个DPP;并引入切线归一化以反映上下文信息。大量实验验证了该方法的有效性。代码已公开于https://github.com/NEC-N-SOGI/msdpp。
原文摘要 · Abstract (English)
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method. Our code is publicly available at https://github.com/NEC-N-SOGI/msdpp.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。