arXiv:2605.07507cs.CLcs.IR2026-05

基于浏览器的LLM系统,零安装提取医学文献结构化信息

TCMIIES: A Browser-Based LLM-Powered Intelligent Information Extraction System for Academic Literature

  • 用可视化界面定义提取模板,无需编程即可配置
  • 在中医文献上准确率超94%,κ值达0.82接近专家水平
  • 本地处理+多模型支持,适合不熟悉技术的研究者使用

学术论文的快速增长催生了从非结构化文本中提取结构化知识的需求。尽管大语言模型(LLMs)具备自然语言理解与信息抽取能力,但现有方案常需专用基础设施、编程技能或微调领域模型,限制了专业研究人员的使用。本文介绍TCMIIES(传统中医信息智能抽取系统),一个基于浏览器、无需安装的平台,利用商用LLM API实现学术文献的结构化信息抽取。系统采用模式引导提示框架,支持自动系统提示生成,用户可通过图形界面自定义提取模式而无需编程。TCMIIES具有纯前端架构,所有信息在浏览器内本地处理,支持五大主流LLM提供商(DeepSeek、OpenAI、Qwen、Zhipu AI及自定义OpenAI兼容接口),实现并发批量处理与自动重试机制,并提供对中文数据库(如CNKI、万方)的智能字段映射。在中医药研究多个场景下的评估显示,结构化输出合规率超过94%,抽取准确率接近但略低于专家间一致性(κ=0.82)。该系统为需要大规模处理文献的领域研究者提供了灵活、隐私保护且成本可控的解决方案。

原文摘要 · Abstract (English)

The rapid growth of academic publications has created a need for tools that extract structured knowledge from unstructured scientific texts. Although large language models (LLMs) can perform natural language understanding and information extraction, existing solutions often require specialized infrastructure, programming expertise, or fine-tuned domain-specific models, which limits their accessibility for researchers in specialized fields. This paper describes TCMIIES (Traditional Chinese Medicine Information Intelligent Extraction System), a browser-based, zero-installation platform that uses commercial LLM APIs to perform structured information extraction from academic literature. The system employs a schema-guided prompting framework with automatic system prompt generation, allowing researchers to define custom extraction schemas through a graphical interface without programming. TCMIIES features a pure front-end architecture that processes all information locally in the browser, supports five major LLM providers (DeepSeek, OpenAI, Qwen, Zhipu AI, and custom OpenAI-compatible endpoints), implements concurrent batch processing with automatic retry mechanisms, and provides intelligent field mapping for Chinese academic databases including CNKI and Wanfang. Evaluation across multiple extraction scenarios in Traditional Chinese Medicine research shows structured output compliance rates exceeding 94\% and extraction accuracy approaching but below expert-level agreement ($κ=0.82$ as reference). The system offers a flexible, privacy-preserving, and cost-effective solution for domain researchers who need to process literature at scale.

信息抽取LLM应用中医文献浏览器工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。