用AI自动完成图书馆主题标引,提升效率并符合专业规范
A Skill-Based AI Agentic Pipeline for Library of Congress Subject Indexing
- 将标引流程拆分为概念分析、筛选、验证和字段生成四步智能技能
- 在10个标题上测试,与专业标引高度一致,仅在细分规则上略有差异
- 适合图书馆自动化系统研发者和信息管理研究者参考
本文提出一种模块化的AI代理技能流水线,用于自动化美国国会图书馆主题标目(LCSH)的标引工作。主题标引涉及分析文献内容、选择受控词表术语,并将其编码为MARC21主题访问字段,是图书馆编目中最耗时的环节之一。该系统将此过程分解为四个顺序执行的代理技能:概念分析、定量筛选、权威性验证和MARC字段合成。每个技能均基于《国会图书馆主题标目手册》(SHM)说明文件和主题分析理论构建领域知识。系统在10个标题上进行评估,数据来自哈佛大学图书馆书目数据库(Alma ILS快照)。结果表明,其概念对齐度与专业标引实践高度一致,但在具体性、细分使用及对2026年国会图书馆政策(以LCGFT 655字段替代形式细分)的遵循上存在差异。
原文摘要 · Abstract (English)
This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject indexing - the process of analyzing a work's aboutness, selecting controlled vocabulary terms, and encoding them as MARC21 subject access fields - is one of the most time-consuming components of library cataloging. The system decomposes this process into four discrete, sequentially executed agent skills: conceptual analysis, quantitative filtering, authority validation, and MARC field synthesis. Each skill encodes domain knowledge drawn directly from Library of Congress Subject Headings Manual (SHM) instruction sheets and subject analysis theory. The pipeline was evaluated against a corpus of ten titles whose existing subject headings were captured from the Harvard Library bibliographic dataset (a snapshot of their Alma ILS). Results demonstrate strong conceptual alignment with professional subject indexing practice, with notable differences in specificity, subdivision practice, and the agent's adherence to the 2026 LC policy discontinuing form subdivisions in favor of LCGFT 655 fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。