arXiv:2412.03817cs.CL2024-12综述被引 5

用跨语言嵌入模型识别健康问卷中的重复问题,提升多语言数据一致性。

Detecting Redundant Health Survey Questions Using Language-agnostic BERT Sentence Embedding (LaBSE)

  • 采用LaBSE模型计算中英文健康问卷语义相似度
  • 跨语言相似度检测准确率超过0.99(AUC)
  • 适合需标准化多语言健康数据的研究者使用

本研究旨在计算公开健康问卷之间的语义相似度,以促进基于调查的个人生成健康数据(PGHD)的标准化。从美国国立卫生研究院共享数据标准库(NIH CDE Repository)、PROMIS、韩国公共健康机构及学术文献中收集了中英文健康问卷问题,涵盖多种健康生活日志领域。通过随机配对生成包含1758个问题对的语义文本相似度(STS)数据集,并由两名专家标注相似度得分。基于该数据集构建了三种分类器:词袋模型、SBERT(基于BERT嵌入)和SBERT-LaBSE(基于跨语言BERT句子嵌入)。采用传统混淆矩阵统计评估算法性能。结果显示,SBERT-LaBSE在跨语言语义相似度判断上表现最佳,其受试者工作特征曲线(ROC)和精确率-召回率曲线下的面积均超过0.99。该模型有效识别了跨语言语义等价问题,但在捕捉细微差异和计算效率方面仍有挑战。未来研究应扩展至更大规模多语言数据集,并优化各健康领域评分的一致性校准与归一化。

原文摘要 · Abstract (English)

The goal of this work was to compute the semantic similarity among publicly available health survey questions in order to facilitate the standardization of survey-based Person-Generated Health Data (PGHD). We compiled various health survey questions authored in both English and Korean from the NIH CDE Repository, PROMIS, Korean public health agencies, and academic publications. Questions were drawn from various health lifelog domains. A randomized question pairing scheme was used to generate a Semantic Text Similarity (STS) dataset consisting of 1758 question pairs. Similarity scores between each question pair were assigned by two human experts. The tagged dataset was then used to build three classifiers featuring: Bag-of-Words, SBERT with BERT-based embeddings, and SBRET with LaBSE embeddings. The algorithms were evaluated using traditional contingency statistics. Among the three algorithms, SBERT-LaBSE demonstrated the highest performance in assessing question similarity across both languages, achieving an Area Under the Receiver Operating Characteristic (ROC) and Precision-Recall Curves of over 0.99. Additionally, it proved effective in identifying cross-lingual semantic similarities.The SBERT-LaBSE algorithm excelled at aligning semantically equivalent sentences across both languages but encountered challenges in capturing subtle nuances and maintaining computational efficiency. Future research should focus on testing with larger multilingual datasets and on calibrating and normalizing scores across the health lifelog domains to improve consistency. This study introduces the SBERT-LaBSE algorithm for calculating semantic similarity across two languages, showing it outperforms BERT-based models and the Bag of Words approach, highlighting its potential to improve semantic interoperability of survey-based PGHD across language barriers.

语义相似度跨语言健康数据问卷优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。