用AI代理协作生成可解释的主题模型,提升透明度与准确性。
Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling

- 多智能体协同完成主题识别、验证与分层,全程可追溯。
- 在BBC数据集上F1达0.95,接近GPT-4.1,优于LDA和BERTopic。
- 适合金融、医疗等需高可解释性的关键场景使用。
Agentopic是一种基于智能体的可解释主题建模新流程,利用大语言模型的推理能力解决传统方法如LDA和BERTopic在主题分配与聚类过程中的不透明问题。该框架通过多个智能体协作完成主题识别、验证、层次化分组及自然语言解释,使用户可追踪主题归属的推理路径,在保持准确性的前提下显著提升可解释性。在英国广播公司(BBC)数据集上,以预设主题为起点,Agentopic实现0.95的F1分数,与GPT-4.1相当,优于LDA的0.93,接近BERTopic的0.98。此外,未预设主题的Agentopic生成了2045个语义连贯的主题,分属六级层次结构,极大丰富了原始五类体系。通过将可解释性嵌入全流程,Agentopic为金融、医疗等关键领域提供透明可靠的替代方案。
原文摘要 · Abstract (English)
Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Existing topic modeling approaches such as Latent Dirichlet Allocation (LDA) and BERTopic often lack transparency on how topics are assigned or grouped. Agentopic addresses this by using multiple agents that collaboratively perform topic identification, validation, hierarchical grouping, and natural language explanation. This design enables users to trace the reasoning behind topic assignments, enhancing interpretability without sacrificing accuracy. When seeded with topics from the British Broadcasting Corporation (BBC) dataset, Agentopic achieves an F1-score of 0.95, matching GPT-4.1, improving on LDA (0.93), and close to BERTopic (0.98). We used Agentopic to augment the BBC dataset with generated explanations to improve the dataset's richness and context. The unseeded Agentopic generated 2045 semantically coherent topics organized across six hierarchical levels, vastly enriching the original five-category structure. By embedding explainability throughout the workflow, Agentopic offers an interpretable alternative to black-box models, making it particularly valuable for crucial applications like finance and healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。