用大模型生成阿尔茨海默病患者行为数据,解决真实数据稀缺问题。
SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients
- 基于大语言模型构建三阶段框架,合成具阿尔茨海默病特征的活动数据。
- 生成数据在人体行为识别任务中提升达79.69%,与真实数据运动特征高度一致。
- 适合智能健康、医疗数据生成领域研究者使用,兼顾隐私与成本。
阿尔茨海默病(AD)已成为全球重大健康挑战,亟需高效的智能健康监测方案。然而,相关行为数据集稀缺严重制约了发展。为此,我们提出 SHADE-AD——一种基于大语言模型(LLM)的、用于生成嵌入AD特征的人体活动数据集的框架。该框架结合公开数据与自收集的99例AD患者数据,合成反映AD特有行为的活动视频。通过三阶段训练机制,扩展了原始采集环境限制下的活动范围。对生成数据的全面评估显示,在人体行为识别(HAR)等下游任务中性能提升最高达79.69%。真实与合成数据间的详细运动指标高度匹配,验证了其真实性与实用性。结果表明,SHADE-AD可为智能健康应用提供低成本、高隐私保护的AD监测解决方案。
原文摘要 · Abstract (English)
Alzheimer's Disease (AD) has become an increasingly critical global health concern, which necessitates effective monitoring solutions in smart health applications. However, the development of such solutions is significantly hindered by the scarcity of AD-specific activity datasets. To address this challenge, we propose SHADE-AD, a Large Language Model (LLM) framework for Synthesizing Human Activity Datasets Embedded with AD features. Leveraging both public datasets and our own collected data from 99 AD patients, SHADE-AD synthesizes human activity videos that specifically represent AD-related behaviors. By employing a three-stage training mechanism, it broadens the range of activities beyond those collected from limited deployment settings. We conducted comprehensive evaluations of the generated dataset, demonstrating significant improvements in downstream tasks such as Human Activity Recognition (HAR) detection, with enhancements of up to 79.69%. Detailed motion metrics between real and synthetic data show strong alignment, validating the realism and utility of the synthesized dataset. These results underscore SHADE-AD's potential to advance smart health applications by providing a cost-effective, privacy-preserving solution for AD monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。