arXiv:2603.08392cs.CL2026-03

用多视角评估框架提升大模型健康建议的可信度与个性化。

COACH meets QUORUM: A Framework and Pipeline for Aligning User, Expert and Developer Perspectives in LLM-generated Health Counselling

  • 构建用户、专家、开发者三方协同的评估框架
  • 实测显示三方可对建议质量达成共识,但对语气等有分歧
  • 适合关注医疗AI可解释性与真实场景落地的研究者

能够收集睡眠、情绪和活动数据的系统,可为慢性病人群提供有价值的生活方式建议。然而这类系统开发极具挑战:不仅要可靠地从用户数据中提取模式,还需结合经验证的医学知识以保障建议质量,并生成符合真实用户情境的内容。本文提出QUORUM,一个统一开发者、专家与用户视角的评估框架,并通过健康日记应用Healthy Chronos(面向癌症患者及幸存者)的实际案例,证明该框架能有效追踪各利益相关方观点的收敛与分歧。同时提出COACH,一种基于大语言模型的个性化生活方式建议生成流水线。应用该框架发现,尽管用户、医学专家与开发人员普遍认可建议的相关性、高质量和可靠性,但在建议语气、对模式识别错误的敏感度以及潜在幻觉问题上存在分歧。研究强调了多利益相关方评估在消费级健康语言技术中的重要性,并展示统一评估框架如何支持可信赖、以患者为中心的NLP系统在真实环境中的部署。

原文摘要 · Abstract (English)

Systems that collect data on sleep, mood, and activities can provide valuable lifestyle counselling to populations affected by chronic disease and its consequences. Such systems are, however, challenging to develop; besides reliably extracting patterns from user-specific data, systems should also contextualise these patterns with validated medical knowledge to ensure the quality of counselling, and generate counselling that is relevant to a real user. We present QUORUM, a new evaluation framework that unifies these developer-, expert-, and user-centric perspectives, and show with a real case study that it meaningfully tracks convergence and divergence in stakeholder perspectives. We also present COACH, a Large Language Model-driven pipeline to generate personalised lifestyle counselling for our Healthy Chronos use case, a diary app for cancer patients and survivors. Applying our framework shows that overall, users, medical experts, and developers converge on the opinion that the generated counselling is relevant, of good quality, and reliable. However, stakeholders also diverge on the tone of the counselling, sensitivity to errors in pattern-extraction, and potential hallucinations. These findings highlight the importance of multi-stakeholder evaluation for consumer health language technologies and illustrate how a unified evaluation framework can support trustworthy, patient-centered NLP systems in real-world settings.

健康AI多视角评估个性化建议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。