用决策理论优化文档筛选停止时机,提升专业搜索效率。
Decision-Theoretic Stopping Rules for Document Screening

- 基于完全信息期望价值设计三种停止策略
- 在专利和医学综述数据集上实现更高净收益
- 适合需权衡成本与回报的法律、医疗检索场景
确定何时停止检索结果审查是多个领域的常见问题。现有技术辅助审查(TAR)中的停止规则旨在达到预设召回率目标,但未考虑审查目的,可能导致次优决策。本文引入决策理论,基于完全信息期望价值推导出三种实用停止策略。该方法应用于专利审查和系统性文献回顾两个专业任务。在CLEF-IP和医学系统性综述数据集上的实验表明,在给定成本与收益设置下,所提方法产生的停止决策更合理,整体净效用优于现有方法。
原文摘要 · Abstract (English)
Deciding when to stop reviewing the results of a search is a common problem with multiple applications. Existing stopping rules developed within Technology-Assisted Review (TAR) aim to achieve a pre-specified recall target and do not take into account the reason for examining the results, potentially leading to sub-optimal recommendations. This paper applies decision theory to the problem and uses it to derive three practical stopping policies based on the Expected Value of Perfect Information. The approach is applied to two professional search tasks: patent examining and systematic reviewing. Experiments on CLEF-IP and medical systematic review datasets show that the proposed approach generally produces more appropriate stopping decisions than existing methods, as demonstrated by higher net utility under the evaluated cost and payoff settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。