通过智能重检索与多模型融合提升生物医学问答精度
Cost-Pragmatic Quality Gating and Selection-Fusion Multi-Model Combiners for BioASQ Phases A+ and B
- 采用双路检索+交叉编码器质量门控,仅对弱支持问题重检索
- 成本降低12%的同时,列表F1和精确率显著提升,验证检索空间充足
- 揭示大模型判别力在选择类任务占优,但融合类任务需额外优化
我们介绍BioASQ Task 14B 2026系统。核心设计包括:在初检薄弱时如何权衡重检索强度,以及如何融合多语言模型答案。检索采用双并行管道——混合初检(密集BGE + BM25 + RRF,BioASQ-13b历史库R@200达99.3%)与代理驱动的PubMed、Europe PMC、iCite分解式检索,并用BGE交叉编码器质量门控识别低置信问题进行选择性重检索。在Task 12B 2024验证集上,成本务实的重检索策略在列表F1和列表精确率上显著优于严格基准,且重检索成本降低12%。固定提示与模型条件下,验证集至测试集13B(不同题集),列表F1绝对提升+0.132,表明检索侧仍有显著提升空间。对于Phase B回答,将多模型集成收益分解为受每题最优基准限制的选择成分与聚合器可超越的融合成分。该分解预测显示,大模型作为裁判在选择主导指标(是/否、多参考ROUGE)上占优,但在融合友好型指标(事实型排名1、列表召回率)的召回成分上存在结构性不足。在Task 13B 2025中,同义词合并解析器在所有头部均取得最高列表召回率,而GPT-5.5单模型仍保持列表F1领先,因其更宽项目集牺牲了精确率。在Task 14B 2026预榜中,我队在八组(阶段x批次)中的三组联合精确值排名第一,赢得四个题型单元格冠军,并在Phase B b3理想模式下位列第一。
原文摘要 · Abstract (English)
We describe our BioASQ Task 14B 2026 system. The work centers on two design decisions: how aggressively to re-retrieve when first-stage retrieval is weak, and how to combine multiple language-model answers. Retrieval unions two parallel pipelines - a hybrid first stage (dense BGE + BM25 + RRF, reaching R@200 = 99.3% on the BioASQ-13b historical archive) and an agent-driven pipeline that decomposes the question over PubMed, Europe PMC, and iCite - with a BGE cross-encoder quality gate flagging weakly-supported questions for selective re-retrieval. On Task 12B 2024 validation, a cost-pragmatic re-retrieval policy beats a skill-strict baseline significantly on list F1 and list precision, at 12% lower re-retrieval cost. Holding prompt and model fixed across val and test 13B (different question sets), list F1 rises by +0.132 absolute on the BioASQ-released gold-input pool, consistent with substantial retrieval-side headroom. For Phase B answering we decompose multi-model ensemble lift into a selection component bounded by the per-question oracle and a fusion component that aggregators can exceed. The decomposition predicts before any experiment that LLM-as-judge wins on selection-dominated metrics (yes/no, multi-reference ROUGE) but is structurally insufficient on the recall component of fusion-friendly metrics (factoid rank-1, list recall). On Task 13B 2025 our synonym-union resolver wins list recall on every head, while GPT-5.5 solo retains the list-F1 lead because the resolver's wider item set costs precision. On the Task 14B 2026 preliminary leaderboard our team places first on the combined-exact aggregate on three of the eight (phase x batch) leaderboards, wins four individual question-type cells, and takes #1 on Phase B b3 ideal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。