arXiv:2502.14748cs.CL2025-02被引 5

LLM看大文档集像找针却找不到,加人管才靠谱

Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models

  • 用人类监督改进LLM生成主题,缓解幻觉和泛化过度
  • 在领域数据上LLM主题太泛,用户难学到有用信息
  • 传统LDA虽不友好但更有效,适合专业探索场景

自然语言处理常用于理解大规模文档集合,近年来从传统主题模型转向大型语言模型(LLM)。本研究评估了无监督、有监督的LLM方法与传统主题模型在两个数据集上的表现。尽管LLM生成的主题更易读且平均胜率更高,但在领域特定数据上产生过于通用的主题,难以帮助用户获取文档知识。加入人类监督可减轻幻觉和泛化问题,但需更高人工成本。相比之下,传统模型如潜在狄利克雷分配(LDA)仍具有效性,但用户体验较差。结果表明,LLM在缺乏人类协助时难以准确描述大规模语料库,尤其在领域数据上受限于上下文长度,面临可扩展性与幻觉挑战。

原文摘要 · Abstract (English)

A common use of NLP is to facilitate the understanding of large document collections, with a shift from using traditional topic models to Large Language Models. Yet the effectiveness of using LLM for large corpus understanding in real-world applications remains under-explored. This study measures the knowledge users acquire with unsupervised, supervised LLM-based exploratory approaches or traditional topic models on two datasets. While LLM-based methods generate more human-readable topics and show higher average win probabilities than traditional models for data exploration, they produce overly generic topics for domain-specific datasets that do not easily allow users to learn much about the documents. Adding human supervision to the LLM generation process improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. In contrast, traditional. models like Latent Dirichlet Allocation (LDA) remain effective for exploration but are less user-friendly. We show that LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints.

主题建模LLM评估人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。