用程序化方法自动评估临床试验偏倚风险,提升准确性和可复现性。
Automated Risk-of-Bias Assessment of Randomized Controlled Trials: A First Look at a GEPA-trained Programmatic Prompting Framework
- 通过代码优化替代人工设计提示词,实现可追踪的自动化评估
- 在7个偏倚领域中整体准确率超人工提示30%-40%,尤其在随机分组上表现突出
- 适合需要高可靠性证据合成的研究者和系统综述团队
评估随机对照试验的偏倚风险对可信证据综合至关重要,但传统方法耗时且评审者间差异大。大型语言模型(LLMs)为自动化提供了可能,但现有方法依赖难以复现、泛化或评估的人工提示设计。本研究提出一种可编程的偏倚评估流程,采用DSPy及其GEPA模块,通过帕累托引导搜索优化LLM推理,并生成可检查的执行轨迹,实现全过程透明可复现。我们在来自7个元分析的100项随机对照试验上评估该方法,应用于开源模型(Mistral Small 3.1与GPT-oss-20b)及商用模型(GPT-5 Nano与GPT-5 Mini)。在报告更清晰的领域(如随机序列生成),GEPA生成的提示表现最优;分配隐藏与受试者盲法方面结果相近,商用模型整体略优。与三个手工提示对比,GEPA在随机序列生成与选择性报告上提升30%-40%准确率,其余领域表现相当甚至更优。结果表明,GEPA能生成一致、可复现的提示,支持在证据合成中结构化、有原则地使用LLMs。
原文摘要 · Abstract (English)
Assessing risk of bias (RoB) in randomized controlled trials is essential for trustworthy evidence synthesis, but the process is resource-intensive and prone to variability across reviewers. Large language models (LLMs) offer a route to automation, but existing methods rely on manually engineered prompts that are difficult to reproduce, generalize, or evaluate. This study introduces a programmable RoB assessment pipeline that replaces ad-hoc prompt design with structured, code-based optimization using DSPy and its GEPA module. GEPA refines LLM reasoning through Pareto-guided search and produces inspectable execution traces, enabling transparent replication of every step in the optimization process. We evaluated the method on 100 RCTs from published meta-analyses across seven RoB domains. GEPA-generated prompts were applied to both open-weight models (Mistral Small 3.1 with GPT-oss-20b) and commercial models (GPT-5 Nano and GPT-5 Mini). In domains with clearer methodological reporting, such as Random Sequence Generation, GEPA-generated prompts performed best, with similar results for Allocation Concealment and Blinding of Participants, while the commercial model performed slightly better overall. We also compared GEPA with three manually designed prompts using Claude 3.5 Sonnet. GEPA achieved the highest overall accuracy and improved performance by 30%-40% in Random Sequence Generation and Selective Reporting, and showed generally comparable, competitively aligned performance in the other domains relative to manual prompts. These findings suggest that GEPA can produce consistent and reproducible prompts for RoB assessment, supporting the structured and principled use of LLMs in evidence synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。