Finder用AI统一检索药学文本图像音频视频,提升搜索精准度。
Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval
- 融合词法与语义向量的混合检索,支持多模态内容统一搜索。
- 处理超29万份文档、3万+视频及1192个音频文件,覆盖98种语言。
- 适合医药研发、监管合规与商业分析人员快速获取跨模态信息。
人工智能正在重塑药学信息检索,传统系统在处理多模态内容和人工标注方面存在瓶颈。Finder是一种可扩展的AI驱动框架,通过混合向量搜索统一文本、图像、音频和视频的检索,结合稀疏词法模型与密集语义模型。其模块化流程可接入多种格式,增强元数据并存储于向量原生后端。系统支持具备推理能力的自然语言搜索,显著提升精度与上下文相关性。目前已处理超过291,400份文档、31,070个视频和1,192个音频文件,涵盖98种语言。通过混合融合、分块处理和元数据感知路由等技术,实现对法规、科研与商业领域的智能信息访问。
原文摘要 · Abstract (English)
AI is transforming pharmaceutical search, where traditional systems struggle with multimodal content and manual curation. Finder is a scalable AI-powered framework that unifies retrieval across text, images, audio, and video using hybrid vector search, combining sparse lexical and dense semantic models. Its modular pipeline ingests diverse formats, enriches metadata, and stores content in a vector-native backend. Finder supports reasoning-aware natural language search, improving precision and contextual relevance. The system has processed over 291,400 documents, 31,070 videos, and 1,192 audio files in 98 languages. Techniques like hybrid fusion, chunking, and metadata-aware routing enable intelligent access across regulatory, research, and commercial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。