arXiv:2504.17753cs.CL2025-04被引 3

对比神经符号系统与ChatGPT在心衰患者饮食咨询中的表现

Conversational Assistants to support Heart Failure Patients: comparing a Neurosymbolic Architecture with ChatGPT

  • 用神经符号架构构建本地对话系统,结合规则与学习
  • 本地系统更准确、任务完成率高且更简洁,但更易出语音错误
  • 适合对准确性要求高的医疗场景,尤其关注可解释性

对话式助手在医疗领域日益普及,尤其得益于大语言模型的发展。然而,仍需针对真实用户开展可控的评估,以揭示传统架构与生成式AI架构的优劣。我们开展了一项组内用户研究,比较了两种版本的对话系统:一种由团队自研的神经符号架构系统,另一种基于ChatGPT。系统功能为帮助心力衰竭患者查询食物含盐量。结果显示,自研系统在准确性、任务完成率和简洁性方面优于ChatGPT;而后者语音错误更少,所需澄清更少。患者对两者无明显偏好。

原文摘要 · Abstract (English)

Conversational assistants are becoming more and more popular, including in healthcare, partly because of the availability and capabilities of Large Language Models. There is a need for controlled, probing evaluations with real stakeholders which can highlight advantages and disadvantages of more traditional architectures and those based on generative AI. We present a within-group user study to compare two versions of a conversational assistant that allows heart failure patients to ask about salt content in food. One version of the system was developed in-house with a neurosymbolic architecture, and one is based on ChatGPT. The evaluation shows that the in-house system is more accurate, completes more tasks and is less verbose than the one based on ChatGPT; on the other hand, the one based on ChatGPT makes fewer speech errors and requires fewer clarifications to complete the task. Patients show no preference for one over the other.

对话系统心衰管理神经符号医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。