arXiv:2604.04071cs.CV2026-04中稿 · CAA 2026 Internati…

用无标签学习发现文化库中的媒体克隆,提升去重效率。

Detecting Media Clones in Cultural Repositories Using a Positive Unlabeled Learning Approach

论文配图:Detecting Media Clones in Cultural Repositories Using a Positive Unlabeled Learning Approach
图 1 · 摘自论文原文
  • 以锚点图像训练轻量编码器,通过相似度评分筛选潜在克隆
  • 在AtticPOT数据集上达F1=90.79,比最优基线提升7.70点
  • 无需负样本,结果可解释,适合人工校对流程

我们将阿提卡陶器档案库(AtticPOT)中馆员参与的重复项发现任务建模为正例-未标记(PU)学习问题。针对每件文物提供一个锚点,训练轻量级查询编码器,利用锚点的增强视图进行学习,并基于潜在空间l_2范数设定可解释阈值对未标记库进行打分。系统生成候选克隆供馆员验证,成功发现此前未被确认的跨记录重复项。在CIFAR-10上取得F1=96.37(AUROC=97.97);在AtticPOT上达到F1=90.79(AUROC=98.99),相比最佳基线(SVDD)在相同轻量主干下提升7.70点。定性“找相似”面板显示视角与条件下邻域稳定。该方法避免显式负样本,提供透明操作点,适用于去重、记录链接及馆员协同工作流。

原文摘要 · Abstract (English)

We formulate curator-in-the-loop duplicate discovery in the AtticPOT repository as a Positive-Unlabeled (PU) learning problem. Given a single anchor per artefact, we train a lightweight per-query Clone Encoder on augmented views of the anchor and score the unlabeled repository with an interpretable threshold on the latent l_2 norm. The system proposes candidates for curator verification, uncovering cross-record duplicates that were not verified a priori. On CIFAR-10 we obtain F1=96.37 (AUROC=97.97); on AtticPOT we reach F1=90.79 (AUROC=98.99), improving F1 by +7.70 points over the best baseline (SVDD) under the same lightweight backbone. Qualitative "find-similar" panels show stable neighbourhoods across viewpoint and condition. The method avoids explicit negatives, offers a transparent operating point, and fits de-duplication, record linkage, and curator-in-the-loop workflows.

去重无监督学习文化遗产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。