arXiv:2604.07562cs.CLcs.AI2026-04ACL被引 1

用大模型推理优化无监督文本聚类,让结果更合理可解释。

Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs

论文配图:Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
图 1 · 摘自论文原文
  • 让大模型当语义裁判,验证并重构聚类结果
  • 在社交媒体数据上提升聚类一致性和标签质量
  • 无需标注即可生成人类认可的聚类标签,适合无监督分析

无监督方法广泛用于从大规模文本中挖掘潜在语义结构,但其输出常包含不连贯、冗余或缺乏依据的聚类,难以在无标签情况下验证。本文提出一种基于推理的聚类优化框架,利用大语言模型(LLMs)作为语义判断者,而非嵌入生成器,对任意无监督聚类算法的输出进行验证与重构。框架包含三个推理阶段:(i) 一致性验证,评估聚类摘要是否由成员文本支持;(ii) 冗余裁决,根据语义重叠合并或剔除候选聚类;(iii) 标签锚定,通过两阶段过程在完全无监督下生成并整合语义相似的标签。该设计将表示学习与结构验证解耦,缓解仅依赖嵌入方法的常见失败模式。我们在两个具有不同互动模式的社交平台真实语料上评估,结果表明该框架在聚类一致性和人类对齐标签质量上持续优于经典主题模型和近期基于表示的方法。人工评估显示与大模型生成标签高度一致,尽管无黄金标准标注。我们还在匹配的时间与数据量条件下进行鲁棒性分析,评估跨平台稳定性。实验表明,大模型推理可作为通用机制,用于验证与优化无监督语义结构,实现无需监督的可靠且可解释的大规模文本分析。

原文摘要 · Abstract (English)

Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redundant, or poorly grounded clusters that are difficult to validate without labeled data. We propose a reasoning-based refinement framework that leverages large language models (LLMs) not as embedding generators, but as semantic judges that validate and restructure the outputs of arbitrary unsupervised clustering algorithms. Our framework introduces three reasoning stages: (i) coherence verification, where LLMs assess whether cluster summaries are supported by their member texts; (ii) redundancy adjudication, where candidate clusters are merged or rejected based on semantic overlap; and (iii) label grounding, where clusters are assigned interpretable labels through a two-stage process that generates and consolidates semantically similar labels in a fully unsupervised manner. This design decouples representation learning from structural validation and mitigates the common failure modes of embedding-only approaches. We evaluate the framework in real-world social media corpora from two platforms with distinct interaction models, demonstrating consistent improvements in cluster coherence and human-aligned labeling quality over classical topic models and recent representation-based baselines. Human evaluation shows strong agreement with LLM-generated labels, despite the absence of gold-standard annotations. We further conduct robustness analysis under matched temporal and volume conditions to assess cross-platform stability. Beyond empirical gains, our results suggest that LLM-based reasoning can serve as a general mechanism for validating and refining unsupervised semantic structure, enabling more reliable and interpretable analysis of large text collections without supervision.

聚类优化大模型推理无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。