arXiv:2502.07352cs.CLcs.AI2025-02中稿 · IRCDL 2025被引 7

用大模型自动评估论文主题模型质量,省去人工标注。

Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation

  • 用大模型分析主题一致性、多样性等关键指标。
  • 在多个数据集上验证,结果稳定且可扩展。
  • 适合需要快速迭代主题模型的研究者使用。

本研究提出一种基于大语言模型(LLMs)的自动化框架,用于评估科学文献中动态演化的主题分类体系。在数字图书馆系统中,主题建模对高效组织与检索学术内容至关重要,帮助研究者穿越复杂的知识图谱。随着研究领域不断增多和演变,传统依赖人工和静态评估的方法难以保持时效性。所提方法利用大模型衡量主题一致性、重复性、多样性及主题-文档匹配度等关键质量维度,无需大量专家标注或狭隘统计指标。通过定制化提示词引导模型评估,确保在不同数据集和建模方法下评估结果的一致性和可解释性。在基准语料库上的实验表明该方法具备鲁棒性、可扩展性与适应性,凸显其作为更全面、动态评估策略的潜力。

原文摘要 · Abstract (English)

This study presents a framework for automated evaluation of dynamically evolving topic taxonomies in scientific literature using Large Language Models (LLMs). In digital library systems, topic modeling plays a crucial role in efficiently organizing and retrieving scholarly content, guiding researchers through complex knowledge landscapes. As research domains proliferate and shift, traditional human centric and static evaluation methods struggle to maintain relevance. The proposed approach harnesses LLMs to measure key quality dimensions, such as coherence, repetitiveness, diversity, and topic-document alignment, without heavy reliance on expert annotators or narrow statistical metrics. Tailored prompts guide LLM assessments, ensuring consistent and interpretable evaluations across various datasets and modeling techniques. Experiments on benchmark corpora demonstrate the method's robustness, scalability, and adaptability, underscoring its value as a more holistic and dynamic alternative to conventional evaluation strategies.

主题模型大模型评估自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。