用大模型自动构建社交媒体事实陈述的分级分类体系。
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media
- 利用大模型生成多粒度主题,构建层级化分类结构。
- 在三个数据集上验证,分类清晰且信息冗余低。
- 适合需要快速梳理海量社交言论的研究者和平台方。
随着社交媒体内容的快速增长,分析网络话语变得愈发复杂。本文提出LLMTaxo,一种借助大语言模型实现社交媒体中事实性陈述自动构建分类体系的新框架,可通过生成多粒度主题构建层次化结构,显著降低信息冗余并提升可访问性。我们还设计了专用的分类体系评估指标,实现全面评估。在三个不同数据集上的实验表明,该框架能生成清晰、连贯且全面的分类体系。在对比模型中,GPT-4o mini 在多数指标上表现最优。该框架灵活性强,对人工干预依赖度低,具有广泛适用潜力。
原文摘要 · Abstract (English)
With the rapid expansion of content on social media platforms, analyzing and comprehending online discourse has become increasingly complex. This paper introduces LLMTaxo, a novel framework leveraging large language models for the automated construction of taxonomies of factual claims from social media by generating topics at multiple levels of granularity. The resulting hierarchical structure significantly reduces redundancy and improves information accessibility. We also propose dedicated taxonomy evaluation metrics to enable comprehensive assessment. Evaluations conducted on three diverse datasets demonstrate LLMTaxo's effectiveness in producing clear, coherent, and comprehensive taxonomies. Among the evaluated models, GPT-4o mini consistently outperforms others across most metrics. The framework's flexibility and low reliance on manual intervention underscore its potential for broad applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。