arXiv:2507.22913cs.CLcs.AI2025-07中稿 · ASIST 2025被引 2

混合模型提升图书主题词生成准确率,避免大模型幻觉。

A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models

  • 用机器学习预测最佳主题词数量,引导大模型生成
  • 对大模型输出进行后处理,纠正为真实主题词表术语
  • 适合图书馆系统、信息检索领域研究者使用

为信息资源提供主题访问是任何图书馆管理系统的核心功能。尽管大语言模型(LLMs)在分类和摘要任务中广泛应用,但其在主题分析方面的潜力尚未充分探索。传统机器学习(ML)模型虽用于多标签分类进行主题分析,但在面对未见案例时表现不佳。大模型虽可替代,但常出现过度生成与幻觉问题。为此,本文提出一种融合嵌入式机器学习模型与大语言模型的混合框架:(1)利用机器学习模型预测最优的美国国会图书馆主题词(LCSH)数量,以指导大模型生成;(2)通过实际LCSH术语对大模型预测结果进行后处理,降低幻觉风险。我们在书籍主题词预测任务上测试了大模型及该混合框架,实验表明,提供初始预测引导并实施后编辑,可获得更受控且术语一致的输出。

原文摘要 · Abstract (English)

Providing subject access to information resources is an essential function of any library management system. Large language models (LLMs) have been widely used in classification and summarization tasks, but their capability to perform subject analysis is underexplored. Multi-label classification with traditional machine learning (ML) models has been used for subject analysis but struggles with unseen cases. LLMs offer an alternative but often over-generate and hallucinate. Therefore, we propose a hybrid framework that integrates embedding-based ML models with LLMs. This approach uses ML models to (1) predict the optimal number of LCSH labels to guide LLM predictions and (2) post-edit the predicted terms with actual LCSH terms to mitigate hallucinations. We experimented with LLMs and the hybrid framework to predict the subject terms of books using the Library of Congress Subject Headings (LCSH). Experiment results show that providing initial predictions to guide LLM generations and imposing post-edits result in more controlled and vocabulary-aligned outputs.

主题分析大模型混合框架信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。