arXiv:2603.23518cs.CLcs.AI2026-03被引 5

用大模型当自动聚类助手,能理解指令还懂自己找分组数。

Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents

  • 让大模型通过推理来理解聚类指令并自主发现数据结构
  • 在28项任务中超越传统嵌入方法和同类模型
  • 适合需要解释性、按指令分组的场景如法律和金融

通用嵌入模型擅长识别语义相似性,但无法遵循用户指令捕捉文本特征。而指令微调的嵌入器虽能对齐指令,却无法自主推断数据潜在结构(如最优聚类数)。为此,我们把指令聚类重构为生成任务,训练大推理模型(LRM)作为自主聚类代理。推理驱动的训练流程使模型能解读高层聚类指令并推断对应隐含分组。为评估该范式,我们引入ReasonCluster基准,包含28个涵盖日常对话、法律案例和财务报告的多样化任务。在多种数据集与聚类场景下实验表明,该方法持续优于强嵌入基线和LRM基线,证明显式推理能实现更忠实且可解释的指令式聚类。

原文摘要 · Abstract (English)

General-purpose embedding models excel at recognizing semantic similarities but fail to capture the characteristics of texts specified by user instructions. In contrast, instruction-tuned embedders can align embeddings with textual instructions yet cannot autonomously infer latent corpus structures, such as determining the optimal number of clusters. To address both limitations, we reframe instruction-following clustering as a generative task and train large reasoning models (LRMs) as autonomous clustering agents. Our reasoning-driven training pipeline enables LRMs to interpret high-level clustering instructions and then infer the corresponding latent groupings. To evaluate this paradigm, we introduce ReasonCluster, a comprehensive benchmark comprising 28 diverse tasks spanning daily dialogue, legal cases, and financial reports. Experiments across diverse datasets and clustering scenarios show that our approach consistently outperforms strong embedding-based methods and LRM baselines, demonstrating that explicit reasoning fosters more faithful and interpretable instruction-based clustering.

聚类大模型指令遵循推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。