arXiv:2411.00041cs.CLcs.AI2024-11

用轻量模型和优化算法,让论文摘要自动分类与答问更准更快。

NeuroSym-BioCAT: Leveraging Neuro-Symbolic Methods for Biomedical Scholarly Document Categorization and Question Answering

  • 结合优化主题模型与精简语言模型,提升文献摘要分类效率。
  • 仅用摘要就能达到高准确率,证明摘要信息足够应对多数生物医学问题。
  • 小模型表现媲美大模型,适合资源有限的医疗信息场景。

生物医学文献摘要数量激增,给高效获取精准信息带来挑战。为此,我们提出一种新方法:将优化的主题建模框架OVB-LDA与BI-POP CMA-ES优化技术结合,用于提升文献摘要分类;同时使用在领域数据上微调的压缩版MiniLM模型进行高精度答案抽取。在三种配置下评估——文献摘要检索、标准文献摘要、标准片段——均优于RYGH和bio-answer finder等现有方法。值得注意的是,仅从文献摘要中提取答案即可实现高准确率,说明摘要对多数生物医学问题已足够。尽管体积小,MiniLM仍表现优异,挑战了‘只有大型模型才能处理复杂任务’的固有观念。结果在多种问题类型和测试批次中均验证了方法的鲁棒性与适应性。未来工作将基于更大规模领域数据优化主题模型,进一步改进MiniLM,并引入大语言模型(LLM)以提升问答的精度与效率。当前仍面临复杂列表型问题处理困难及评估指标不一致等挑战。

原文摘要 · Abstract (English)

The growing volume of biomedical scholarly document abstracts presents an increasing challenge in efficiently retrieving accurate and relevant information. To address this, we introduce a novel approach that integrates an optimized topic modelling framework, OVB-LDA, with the BI-POP CMA-ES optimization technique for enhanced scholarly document abstract categorization. Complementing this, we employ the distilled MiniLM model, fine-tuned on domain-specific data, for high-precision answer extraction. Our approach is evaluated across three configurations: scholarly document abstract retrieval, gold-standard scholarly documents abstract, and gold-standard snippets, consistently outperforming established methods such as RYGH and bio-answer finder. Notably, we demonstrate that extracting answers from scholarly documents abstracts alone can yield high accuracy, underscoring the sufficiency of abstracts for many biomedical queries. Despite its compact size, MiniLM exhibits competitive performance, challenging the prevailing notion that only large, resource-intensive models can handle such complex tasks. Our results, validated across various question types and evaluation batches, highlight the robustness and adaptability of our method in real-world biomedical applications. While our approach shows promise, we identify challenges in handling complex list-type questions and inconsistencies in evaluation metrics. Future work will focus on refining the topic model with more extensive domain-specific datasets, further optimizing MiniLM and utilizing large language models (LLM) to improve both precision and efficiency in biomedical question answering.

文献分类问答系统小模型生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。