arXiv:2601.07684cs.IR2026-01

AptaFind让研究人员1小时处理900个适配体目标,自动从文献中提取序列或推荐优质参考文献。

AptaFind: A lightweight local interface for automated aptamer curation from scientific literature

  • 结合本地大模型与确定性算法,离线运行不依赖云端
  • 在300个目标上实现79%直接序列提取、84%找到研究线索
  • 适合实验室人员快速获取文献支持,尤其擅长文献筛选

适配体研究者面临文献分散在各类出版物、补充材料和数据库中的难题,每次检索耗时数小时。AptaFind通过三层智能架构解决这一问题,将科研挖掘视为连续谱而非二元成败。系统在可操作时直接提取序列,失败时提供经筛选的研究线索,必要时进行全文献发现以增强可信度。该系统结合本地语言模型进行语义理解,搭配确定性算法确保可靠性,无需云服务或订阅。在德克萨斯大学适配体数据库的300个目标上验证:84%的目标找到了相关文献,84%获得经筛选的研究线索,79%成功提取直接序列,单机每小时可处理超过900个目标。结果表明,即使无法直接提取序列,自动化仍能通过快速定位高质量参考文献,提供可行动的科研情报。

原文摘要 · Abstract (English)

Aptamer researchers face a literature landscape scattered across publications, supplements, and databases, with each search consuming hours that could be spent at the bench. AptaFind transforms this navigation problem through a three-tier intelligence architecture that recognizes research mining is a spectrum, not a binary success or failure. The system delivers direct sequence extraction when possible, curated research leads when extraction fails, and exhaustive literature discovery for additional confidence. By combining local language models for semantic understanding with deterministic algorithms for reliability, AptaFind operates without cloud dependencies or subscription barriers. Validation across 300 University of Texas Aptamer Database targets demonstrates 84 % with some literature found, 84 % with curated research leads, and 79 % with a direct sequence extraction, at a laptop-compute rate of over 900 targets an hour. The platform proves that even when direct sequence extraction fails, automation can still deliver the actionable intelligence researchers need by rapidly narrowing the search to high quality references.

文献挖掘适配体本地推理自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。