arXiv:2506.19268cs.HCcs.CR2025-06

构建首个大规模医疗应用隐私与信任评论数据集,助力患者隐私研究。

Health App Reviews for Privacy & Trust (HARPT): A Corpus for Analyzing Patient Privacy Concerns, Trust in Providers and Trust in Applications

  • 通过多阶段方法构建含48万条标注评论的HARPT数据集
  • 涵盖应用信任、医生信任和隐私担忧共7类标签,覆盖率达95%以上
  • 适合医疗信息学、可用隐私研究者使用,支持可复现分析

背景:远程医疗和患者门户类移动应用(统称eHealth应用)的用户评论是未经筛选的患者反馈重要来源,揭示了患者对隐私与信任的真实看法。但缺乏大规模、带标注的隐私与信任专用数据集,限制了自然语言处理技术在该领域的系统性分析。目标:本文旨在构建并基准测试健康应用评论中的隐私与信任(HARPT)数据集,推动患者隐私与信任研究。方法:采用多阶段数据构建策略,整合关键词过滤、迭代式人工标注、针对性数据增强及基于Transformer的弱监督分类。从中筛选出7,000条评论进行人工标注,用于模型开发与评估。最终数据集用于基准测试多种机器学习模型。结果:HARPT包含48万条患者评论,覆盖7个类别,涵盖应用信任(TA)、医生信任(TP)和隐私担忧(PC)等关键维度。提供了多种模型在人工标注子集上的全面基准性能,建立了未来研究基线。结论:HARPT是推进eHealth领域隐私与信任研究的重要资源。通过提供大规模、标注数据集与初步基准,本工作支持可用隐私与信任研究的可复现性。HARPT已开源发布。

原文摘要 · Abstract (English)

Background: User reviews of Telehealth and Patient Portal mobile applications (apps) hereon referred to as electronic health (eHealth) apps are a rich source of unsolicited patient feedback, revealing critical insights into patient perceptions. However, the lack of large-scale, annotated datasets specific to privacy and trust has limited the ability of researchers to systematically analyze these concerns using natural language processing (NLP) techniques. Objective: This study aims to develop and benchmark Health App Reviews for Privacy & Trust (HARPT), a large-scale annotated corpus of patient reviews from eHealth apps to advance research in patient privacy and trust. Methods: We employed a multistage data construction strategy. This integrated keyword-based filtering, iterative manual labeling with review, targeted data augmentation, and weak supervision using transformer-based classifiers. A curated subset of 7,000 reviews was manually annotated to support machine learning model development and evaluation. The resulting dataset was used to benchmark a broad range of models. Results: The HARPT corpus comprises 480,000 patient reviews annotated across seven categories capturing critical aspects of trust in the application (TA), trust in the provider (TP), and privacy concerns (PC). We provide comprehensive benchmark performance for a range of machine learning models on the manually annotated subset, establishing a baseline for future research. Conclusions: The HARPT corpus is a significant resource for advancing the study of privacy and trust in the eHealth domain. By providing a large-scale, annotated dataset and initial benchmarks, this work supports reproducible research in usable privacy and trust within health informatics. HARPT is released under an open resource license.

医疗健康隐私分析评论数据集信任建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。