arXiv:2606.24094cs.CV2026-06

用文本指引统一图像聚类,无需针对每种场景重新训练。

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

论文配图:Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
图 1 · 摘自论文原文
  • 通过生成概念代理实现文本指引下的嵌入表示
  • 基于最小生成树的LLM遍历提升复杂语义判断效率
  • 支持从通用到细粒度、长尾分布等多样聚类场景

跨不同聚类场景的统一图像聚类仍面临根本性挑战。本文提出首个通用框架——指南驱动图像聚类代理,通过文本指南弥合任务间差距。为在不进行任务特异性训练的前提下融入复杂指南,我们提出生成式概念代理建模,通过概念代理提取生成具备指南感知能力的嵌入。对于需自动发现聚类的场景,引入基于最小生成树的LLM遍历机制,仅对复杂语义判断应用LLM推理。该框架可泛化至涵盖从通用到细粒度分类、从全局到局部标准、从均衡到长尾分布的多样化聚类任务,在多种聚类任务中持续优于专用方法。

原文摘要 · Abstract (English)

Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the first universal framework that bridges these gaps through textual guidelines. To incorporate complex guidelines without task-specific training, we propose Generative Concept Proxy Modeling, which generates guideline-aware embeddings via concept proxy extraction. For scenarios requiring automatic cluster discovery, we introduce LLM Traversal based on Minimum Spanning Tree that selectively applies LLM reasoning for complex semantic judgments. Our method generalizes across diverse clustering scenarios spanning from general to fine-grained categorization, from global to local criteria, and from balanced to long-tail distributions. Our framework consistently outperforms specialized methods across diverse clustering tasks.

图像聚类大模型文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。