用主题模型+大模型实现农业文本的可解释检索
AgriLens: Semantic Retrieval in Agricultural Texts Using Topic Modeling and Language Models
- 结合BERTopic与语言模型生成主题标签和摘要
- 支持零样本主题标注与向量搜索的语义检索
- 适合缺乏标注数据的农业知识管理场景
随着非结构化文本数量持续增长,亟需可扩展且可解释的方法来组织、摘要和检索信息。本文提出一个统一框架,用于农业文本的可解释主题建模、零样本主题标注及主题引导的语义检索。基于BERTopic提取语义一致的主题,并将其转化为结构化提示,使语言模型能以零样本方式生成有意义的主题标签和摘要。通过密集嵌入与向量搜索支持查询与文档探索,专用评估模块则用于衡量主题连贯性与偏差。该框架在标注数据有限的专精领域中实现了可扩展的可解释信息访问。
原文摘要 · Abstract (English)
As the volume of unstructured text continues to grow across domains, there is an urgent need for scalable methods that enable interpretable organization, summarization, and retrieval of information. This work presents a unified framework for interpretable topic modeling, zero-shot topic labeling, and topic-guided semantic retrieval over large agricultural text corpora. Leveraging BERTopic, we extract semantically coherent topics. Each topic is converted into a structured prompt, enabling a language model to generate meaningful topic labels and summaries in a zero-shot manner. Querying and document exploration are supported via dense embeddings and vector search, while a dedicated evaluation module assesses topical coherence and bias. This framework supports scalable and interpretable information access in specialized domains where labeled data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。