arXiv:2603.26592cs.LGcs.AI2026-03中稿 · publication in Com…被引 1

用二维可视化辅助选样,提升生物医学时序数据标注效率与质量

Evaluating Interactive 2D Visualization as a Sample Selection Strategy for Biomedical Time-Series Data Annotation

  • 通过交互式2D可视化筛选样本,提升标注效率
  • 在多类任务中,2D可视化显著优于随机和远点优先采样法
  • 适合有经验标注者,且能提升标注过程的趣味性

生物医学机器学习模型依赖准确标签,但时序数据标注仍具挑战。本研究比较了三种样本选择方法:随机采样(RND)、远点优先遍历(FAFT)以及基于交互式2D可视化(2DVs)的探索方法。在婴儿运动评估(IMA)和语音情绪识别(SER)四个分类任务中,12名标注员(专家与非专家)在有限预算下进行标注,并开展事后实验评估。结果表明,2DVs在跨标注者标签聚合中表现最佳;在IMA任务中有效捕捉稀有类别,但因标注预算有限导致标注者间一致性下降,个体标注者训练模型时性能反而不如FAFT;在SER任务中,2DVs对专家更优,对非专家亦可达到相似效果。风险分析显示,当标注者数量或水平不确定时,RND最安全,而2DVs因标签分布差异大风险最高。访谈发现,2DVs使标注过程更具吸引力。总体而言,2DVs在预算不严苛时是极具前景的标注策略。

原文摘要 · Abstract (English)

Reliable machine-learning models in biomedical settings depend on accurate labels, yet annotating biomedical time-series data remains challenging. Algorithmic sample selection may support annotation, but evidence from studies involving real human annotators is scarce. Consequently, we compare three sample selection methods for annotation: random sampling (RND), farthest-first traversal (FAFT), and a graphical user interface-based method enabling exploration of complementary 2D visualizations (2DVs) of high-dimensional data. We evaluated the methods across four classification tasks in infant motility assessment (IMA) and speech emotion recognition (SER). Twelve annotators, categorized as experts or non-experts, performed data annotation under a limited annotation budget, and post-annotation experiments were conducted to evaluate the sampling methods. Across all classification tasks, 2DV performed best when aggregating labels across annotators. In IMA, 2DV most effectively captured rare classes, but also exhibited greater annotator-to-annotator label distribution variability resulting from the limited annotation budget, decreasing classification performance when models were trained on individual annotators' labels; in these cases, FAFT excelled. For SER, 2DV outperformed the other methods among expert annotators and matched their performance for non-experts in the individual-annotator setting. A failure risk analysis revealed that RND was the safest choice when annotator count or annotator expertise was uncertain, whereas 2DV had the highest risk due to its greater label distribution variability. Furthermore, post-experiment interviews indicated that 2DV made the annotation task more interesting and enjoyable. Overall, 2DV-based sampling appears promising for biomedical time-series data annotation, particularly when the annotation budget is not highly constrained.

数据标注时序数据可视化生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。