用大模型自动打电话做医疗调查,准确率达98%。
Automated Survey Collection with LLM-based Conversational Agents
- 用大模型当电话调查员,自动拨号提问并记录对话。
- 对话转录错误率7.7%,但答案提取准确率达98%。
- 适合想低成本大规模收集医疗数据的研究者。
传统电话调查虽广泛用于生物医学和医疗数据采集,但成本高、人力密集且难以规模化。为此,我们提出一种基于对话式大语言模型(LLM)的端到端调查收集框架。该框架包括研究者设计问卷并招募参与者,由大模型驱动的电话代理自动拨号并执行调查,另一大模型(GPT-4o)分析调查过程中的对话转录文本,并将结果存入数据库。为验证框架有效性,我们招募了8名参与者(5名母语者,3名非母语者),完成了40次调查。评估结果显示,尽管对话转录的平均每行词错误率为7.7%,但GPT-4o仍能以平均98%的准确率提取调查回答。参与者虽指出大模型偶有失误,但普遍认为其能清晰传达调查目的,理解力良好,互动自然流畅。研究证明,该方法可显著降低人工负担,为真实世界中全自动医疗电话调查系统提供可行路径。
原文摘要 · Abstract (English)
Objective: Traditional phone-based surveys are among the most accessible and widely used methods to collect biomedical and healthcare data, however, they are often costly, labor intensive, and difficult to scale effectively. To overcome these limitations, we propose an end-to-end survey collection framework driven by conversational Large Language Models (LLMs). Materials and Methods: Our framework consists of a researcher responsible for designing the survey and recruiting participants, a conversational phone agent powered by an LLM that calls participants and administers the survey, a second LLM (GPT-4o) that analyzes the conversation transcripts generated during the surveys, and a database for storing and organizing the results. To test our framework, we recruited 8 participants consisting of 5 native and 3 non-native english speakers and administered 40 surveys. We evaluated the correctness of LLM-generated conversation transcripts, accuracy of survey responses inferred by GPT-4o and overall participant experience. Results: Survey responses were successfully extracted by GPT-4o from conversation transcripts with an average accuracy of 98% despite transcripts exhibiting an average per-line word error rate of 7.7%. While participants noted occasional errors made by the conversational LLM agent, they reported that the agent effectively conveyed the purpose of the survey, demonstrated good comprehension, and maintained an engaging interaction. Conclusions: Our study highlights the potential of LLM agents in conducting and analyzing phone surveys for healthcare applications. By reducing the workload on human interviewers and offering a scalable solution, this approach paves the way for real-world, end-to-end AI-powered phone survey collection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。