让图像标签同时支持固定类别和用户自定义,实现更灵活的多模态识别。
OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
- 用视觉与文本联合表示,动态融合固定标签和开放标签。
- 在两个数据集上整体F1达0.81和0.75,开放标签F1高达0.99和0.97。
- 适合需要灵活标签扩展的图文应用,如社交媒体内容管理。
我们提出OTTER,一种统一的开集多标签打标框架,兼顾预定义类别集的稳定性与用户自定义标签的灵活性。OTTER基于大规模分层组织的多模态数据集构建,数据来自多种在线来源,通过自动化视觉-语言标注与人工修正相结合的混合流程标注。利用多头注意力架构,OTTER联合对齐视觉与文本表示,同时匹配固定标签和开放标签嵌入,实现动态且语义一致的打标。在两个基准数据集上,OTTER持续优于对比基线:在Otter数据集上整体F1达0.81,在Favorite数据集上为0.75,分别领先第二名0.10和0.02。对于开放标签,其表现接近完美,F1分别为0.99和0.97,同时在预定义标签上保持竞争力。结果表明,OTTER有效实现了闭集一致性与开集灵活性的融合,适用于多模态打标任务。
原文摘要 · Abstract (English)
We introduce OTTER, a unified open-set multi-label tagging framework that harmonizes the stability of a curated, predefined category set with the adaptability of user-driven open tags. OTTER is built upon a large-scale, hierarchically organized multi-modal dataset, collected from diverse online repositories and annotated through a hybrid pipeline combining automated vision-language labeling with human refinement. By leveraging a multi-head attention architecture, OTTER jointly aligns visual and textual representations with both fixed and open-set label embeddings, enabling dynamic and semantically consistent tagging. OTTER consistently outperforms competitive baselines on two benchmark datasets: it achieves an overall F1 score of 0.81 on Otter and 0.75 on Favorite, surpassing the next-best results by margins of 0.10 and 0.02, respectively. OTTER attains near-perfect performance on open-set labels, with F1 of 0.99 on Otter and 0.97 on Favorite, while maintaining competitive accuracy on predefined labels. These results demonstrate OTTER's effectiveness in bridging closed-set consistency with open-vocabulary flexibility for multi-modal tagging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。