arXiv:2506.08486cs.AI2025-06被引 4

构建可信赖的健康数字孪生系统,提升AI医疗助手的安全与透明性。

RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being

  • 设计结构化提示引擎,动态提取输入槽位,减少幻觉风险。
  • 在四大健康场景中实现90%以上伦理合规与指令遵循,性能领先。
  • 适合关注AI医疗安全、可解释性的研究者与开发者使用。

大语言模型(LLM)的发展为健康数字孪生带来了新可能,但其在消费级健康应用中存在幻觉、偏见、透明度不足及伦理滥用等风险。针对世卫组织等机构建议,我们提出负责任健康数字孪生(RHealthTwin)框架,支持多模态输入,通过健康导向的LLM生成安全、相关且可解释的响应。核心是负责任提示引擎(RPE),它动态提取预定义槽位以结构化输入,替代传统非结构化提示,降低幻觉风险。该引擎使模型输出更具上下文感知、个性化、公平可靠且可解释。系统通过用户反馈循环持续优化提示结构。我们在心理健康支持、症状分诊、营养规划和活动指导四个领域评估,RPE在基准数据集上达到BLEU=0.41、ROUGE-L=0.63、BERTScore=0.89;LLM-as-judge评估显示伦理合规与指令遵循超过90%,优于基线策略。RHealthTwin为负责任的基于LLM的健康应用提供前瞻性基础。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has created new possibilities for digital twins in healthcare. However, the deployment of such systems in consumer health contexts raises significant concerns related to hallucination, bias, lack of transparency, and ethical misuse. In response to recommendations from health authorities such as the World Health Organization (WHO), we propose Responsible Health Twin (RHealthTwin), a principled framework for building and governing AI-powered digital twins for well-being assistance. RHealthTwin processes multimodal inputs that guide a health-focused LLM to produce safe, relevant, and explainable responses. At the core of RHealthTwin is the Responsible Prompt Engine (RPE), which addresses the limitations of traditional LLM configuration. Conventionally, users input unstructured prompt and the system instruction to configure the LLM, which increases the risk of hallucination. In contrast, RPE extracts predefined slots dynamically to structure both inputs. This guides the language model to generate responses that are context aware, personalized, fair, reliable, and explainable for well-being assistance. The framework further adapts over time through a feedback loop that updates the prompt structure based on user satisfaction. We evaluate RHealthTwin across four consumer health domains including mental support, symptom triage, nutrition planning, and activity coaching. RPE achieves state-of-the-art results with BLEU = 0.41, ROUGE-L = 0.63, and BERTScore = 0.89 on benchmark datasets. Also, we achieve over 90% in ethical compliance and instruction-following metrics using LLM-as-judge evaluation, outperforming baseline strategies. We envision RHealthTwin as a forward-looking foundation for responsible LLM-based applications in health and well-being.

数字孪生AI医疗可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。