用大模型自动分析医疗模拟对话,兼顾准确、快速与环保。
Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact
- 设计不同提示词和批量处理策略,优化分析效率。
- 批量越大越快越省电,但准确率下降。
- 适合需要实时反馈的医疗培训场景。
研究表明,对话在团队构建共同理解、协调行动和塑造学习成果中起核心作用。分析对话内容是推进团队学习理论和设计协作学习环境的关键,但以往依赖人工定性编码,耗时费力。大语言模型(LLM)为自动化提升多模态学习分析中的对话层提供了新可能,近期研究显示其可通过少量示例提示接近人工编码效果。然而,现有工作多关注复现人类编码准确性,未解决更关键的教育问题:如何设计提示,使LLM能在真实场景(如现场医疗模拟)中快速准确完成对话标注,且兼顾响应速度、计算成本与可持续性?本文研究提示设计与批量策略对团队医疗模拟复盘中编码准确性、处理时间和环境影响的平衡。基于包含11,647个发言、6种对话构念的语料库,对比4种提示设计在不同批量下的表现,评估编码性能、处理时间、能耗及各项指标间的权衡。结果表明,增加批量可提升速度并降低能耗,但会损害编码性能。本研究不仅证明了基于LLM的定性分析可行性,还为在时效性、隐私性和可持续性要求高的场景中规模化对话分析提供实用指导。
原文摘要 · Abstract (English)
Research shows that dialogue, the interactive process through which participants articulate their thinking, plays a central role in constructing shared understanding, coordinating action, and shaping learning outcomes in teams. Analysing dialogue content has been central to advancing team learning theory and informing the design of computer-supported collaborative learning environments, yet this progress has depended on labour-intensive qualitative coding. LLMs offer new possibilities for automating and enhancing the dialogue layer within emerging multimodal learning analytics approaches, with recent studies showing that they can approximate human coding through few-shot prompting. However, prior work has focused on replicating human coding accuracy for research purposes, rather than addressing a more educationally consequential question: how can we design prompts that allow an LLM to label team dialogue accurately and fast enough to be useful in real settings, such as in-person healthcare simulations, where results must be returned quickly and computational cost and sustainability also matter? This paper investigates how prompt design and batching strategies can be optimised to balance coding accuracy, processing time, and environmental impact in team-based healthcare simulation debriefing. Using a dataset of 11,647 utterances coded across 6 dialogue constructs, we compared 4 prompt designs across varying batch sizes, evaluating coding performance, processing time, and energy consumption, as well as the trade-offs between these metrics. Results indicate that increasing batch size improves speed and reduces energy use, but negatively impacts coding performance. Beyond demonstrating the feasibility of LLM-based qualitative analysis, this study offers practical guidance for scaling dialogue analytics in contexts where timeliness, privacy, and sustainability are critical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。