arXiv:2507.21103cs.IRcs.CL2025-07

用大模型+检索增强生成,自动分析药品说明书。

Analise Semantica Automatizada com LLM e RAG para Bulas Farmaceuticas

  • 结合向量检索与大模型生成,实现文档语义解析。
  • 在药品说明书上准确率高,响应速度快且一致。
  • 适合医药、医疗信息化领域快速处理非结构化文本。

数字文档在学术、商业和医疗领域的生成速度持续增长,给非结构化信息的高效提取与分析带来新挑战。本文研究将检索增强生成(RAG)架构与大规模语言模型(LLMs)结合,用于自动化分析PDF格式文档。该方法整合嵌入向量搜索、语义数据提取及上下文相关自然语言生成技术。通过从官方公开渠道获取的药品说明书进行实验验证,采用准确率、完整性、响应速度和一致性等指标评估语义查询效果。结果表明,RAG与LLM的结合显著提升了对非结构化技术文本的智能信息检索与理解能力。

原文摘要 · Abstract (English)

The production of digital documents has been growing rapidly in academic, business, and health environments, presenting new challenges in the efficient extraction and analysis of unstructured information. This work investigates the use of RAG (Retrieval-Augmented Generation) architectures combined with Large-Scale Language Models (LLMs) to automate the analysis of documents in PDF format. The proposal integrates vector search techniques by embeddings, semantic data extraction and generation of contextualized natural language responses. To validate the approach, we conducted experiments with drug package inserts extracted from official public sources. The semantic queries applied were evaluated by metrics such as accuracy, completeness, response speed and consistency. The results indicate that the combination of RAG with LLMs offers significant gains in intelligent information retrieval and interpretation of unstructured technical texts.

大模型RAG药品说明书语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。