用不确定性感知的智能选对法,减少人工排序工作量11%~16%
Dodgersort: Uncertainty-Aware VLM-Guided Human-in-the-Loop Pairwise Ranking
- 基于视觉语言模型预排序+概率集成,动态选最有信息量的图像对
- 医学、历史、审美任务中减少11%~16%标注量,可靠性提升
- 可提取5~20倍于基线的排序信息,适合高精度高效标注场景
成对比较标注因其更高的评分者间一致性正逐渐兴起,但全部比较需二次方成本。我们提出Dodgersort,利用基于CLIP的分层预排序、神经排序头与概率集成(Elo、BTL、GP)、认知与随机不确定性分解,以及基于信息论的配对选择机制,在减少人工比较的同时提升排序可靠性。在医学影像、历史年代判定和美学评价任务中,该方法实现11%~16%的标注量缩减并提升评分者间一致性。跨数据集四组消融实验表明,神经适应与集成不确定性是性能提升的关键。在具有真实年龄标签的FG-NET数据集上,该框架每轮比较提取的排名信息为基线的5~20倍,实现了准确率-效率的帕累托最优平衡。
原文摘要 · Abstract (English)
Pairwise comparison labeling is emerging as it yields higher inter-rater reliability than conventional classification labeling, but exhaustive comparisons require quadratic cost. We propose Dodgersort, which leverages CLIP-based hierarchical pre-ordering, a neural ranking head and probabilistic ensemble (Elo, BTL, GP), epistemic--aleatoric uncertainty decomposition, and information-theoretic pair selection. It reduces human comparisons while improving the reliability of the rankings. In visual ranking tasks in medical imaging, historical dating, and aesthetics, Dodgersort achieves a 11--16\% annotation reduction while improving inter-rater reliability. Cross-domain ablations across four datasets show that neural adaptation and ensemble uncertainty are key to this gain. In FG-NET with ground-truth ages, the framework extracts 5--20$\times$ more ranking information per comparison than baselines, yielding Pareto-optimal accuracy--efficiency trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。