用开源大模型自动识别文本中的隐含概念并分类,提升分析准确性。
Concept Navigation and Classification via Open-Source Large Language Model Processing
- 结合自动摘要与人工校验,迭代优化概念识别
- 在多类文本数据中实现高精度框架与话题分类
- 适合政策分析、媒体研究等需要深度理解文本的场景
本文提出一种新型方法框架,利用开源大型语言模型(LLMs)从文本数据中检测和分类潜在构念,包括框架、叙事和主题。该混合方法结合自动化摘要与人机协同验证,通过迭代采样和专家精修,确保方法论稳健性和概念精确性。在人工智能政策辩论、加密新闻报道及20 Newsgroups数据集等多样数据集上应用,展现了在复杂政治话语、媒体框架与主题分类任务中的系统性分析能力。
原文摘要 · Abstract (English)
This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The proposed hybrid approach combines automated summarization with human-in-the-loop validation to enhance the accuracy and interpretability of construct identification. By employing iterative sampling coupled with expert refinement, the framework guarantees methodological robustness and ensures conceptual precision. Applied to diverse data sets, including AI policy debates, newspaper articles on encryption, and the 20 Newsgroups data set, this approach demonstrates its versatility in systematically analyzing complex political discourses, media framing, and topic classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。