arXiv:2506.21625cs.CLcs.AI2025-06被引 2

提出Doc2SAR框架,高效精准提取文献中的药物结构-活性关系

Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents

  • 融合领域工具与微调大模型,协同提升文档解析能力
  • 在200篇文献上达到80.78%表格召回率,远超GPT-4o的29.3%
  • 适合药物研发人员快速从论文中挖掘关键化合物数据

从科学文献和专利中提取分子结构-活性关系(SAR)对药物发现和材料研究至关重要。然而,由于文档格式多样及现有方法局限,该任务仍具挑战:基于规则的方法依赖固定模板,泛化能力差;通用多模态大模型在版面检测和光学化学结构识别(OCSR)等专业任务上准确率不足。为此,我们构建了专用于评估SAR提取方法的严格标注基准DocSAR-200(200篇科学文献)。同时提出Doc2SAR框架,通过监督微调(SFT)增强的多模态大模型与领域专用工具协同工作。大量实验表明,Doc2SAR在多种文档类型上均达领先性能,整体表格召回率达80.78%,较端到端GPT-4o高出51.48%。此外,该框架推理高效,配套提供网页应用,具备实际可用性。

原文摘要 · Abstract (English)

Extracting molecular structure-activity relationships (SARs) from scientific literature and patents is essential for drug discovery and materials research. However, this task remains challenging due to heterogeneous document formats and limitations of existing methods. Specifically, rule-based approaches relying on rigid templates fail to generalize across diverse document layouts, while general-purpose multimodal large language models (MLLMs) lack sufficient accuracy and reliability for specialized tasks, such as layout detection and optical chemical structure recognition (OCSR). To address these challenges, we introduce DocSAR-200, a rigorously annotated benchmark of 200 scientific documents designed specifically for evaluating SAR extraction methods. Additionally, we propose Doc2SAR, a novel synergistic framework that integrates domain-specific tools with MLLMs enhanced via supervised fine-tuning (SFT). Extensive experiments demonstrate that Doc2SAR achieves state-of-the-art performance across various document types, significantly outperforming leading end-to-end baselines. Specifically, Doc2SAR attains an overall Table Recall of 80.78% on DocSAR-200, exceeding end2end GPT-4o by 51.48%. Furthermore, Doc2SAR demonstrates practical usability through efficient inference and is accompanied by a web app.

药物发现结构-活性关系多模态模型知识提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。