用混合模型自动识别和分组毛绒装照片中的角色身份
Fursee: Hybrid YOLO-DINOv3 Framework for Fursuit Identity Retrieval and Clustering

- YOLO定位头部区域,DINOv3结合弧度损失增强特征区分度
- 自动优化聚类参数,比主流多模态模型在各项指标上更优
- 专为毛绒装设计数据集与流程,适合爱好者社群管理应用
全球毛绒装大会产生大量毛绒装照片,人工整理成本高昂,亟需自动化身份检索与聚类方案。现有通用多模态模型缺乏对复杂毛绒装场景的优化,且该任务无公开基准数据集。为此,我们构建了专用毛绒装图像数据集,并提出三阶段混合框架Fursee实现身份检索与聚类。首先,使用YOLO检测并裁剪高分辨率毛绒装头部区域,提升小目标与重叠目标的定位精度;其次,采用ArcFace优化DINOv3特征嵌入,在特征超球面上扩大不同身份间的角间距;最后,通过轮廓系数驱动搜索自动选择DBSCAN最优超参数,避免手动设定半径。检索与聚类实验表明,该方法在所有评估指标上均优于GPT5.5、Claude Opus 4.8及Qwen3.7-Plus等主流多模态模型,实现了对毛绒装头部的高效检索与分组性能。
原文摘要 · Abstract (English)
Global furry conventions produce massive fursuit photographs, while manual sorting brings heavy labor costs and calls for automatic identity retrieval and clustering solutions. General multimodal models lack dedicated optimization for complex fursuit scenes, and no public benchmark dataset exists for this task. To fill this gap, we build a specialized fursuit image dataset and present a three-stage hybrid pipeline Fursee for fursuit identity retrieval and clustering. First, YOLO detects and crops high-resolution fursuit head patches to improve localization of small and overlapping targets. Second, ArcFace optimizes DINOv3 embeddings to enlarge angular separation between different identities on the feature hypersphere. Third, DBSCAN performs unsupervised clustering, with silhouette-coefficient-driven search automatically selecting optimal hyperparameters rather than fixed manual radius. Retrieval and clustering experiments verify that our pipeline outperforms mainstream multimodal models including GPT5.5, Claude Opus 4.8 and Qwen3.7-Plus on all evaluation metrics, achieving competitive performance for fursuit head retrieval and grouping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。