arXiv:2604.06138cs.SDcs.AI2026-04被引 2

构建合成医患对话数据集,用于长时音频摘要任务。

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

  • 用角色驱动生成对话,模拟真实医患交流
  • 生成8800段对话,含1300小时音频与参考病历
  • 适合研究长序列语音理解与医疗自动化系统

长时音频推理在训练数据和评估方面均严重不足。现有基准多针对短上下文任务,而最相关的开放式生成任务因自动评估困难难以开展。本文提出一种合成数据生成流程,既可作为训练资源,也可作为受控评估环境,并应用于首次就诊医患对话,以生成SOAP病历为任务目标。该流程包含三个阶段:基于角色的对话生成、带重叠/停顿建模的多人语音合成及室内外声学与音效模拟、基于大模型的参考SOAP病历生成,全部基于开源权重模型构建。我们发布了8800段合成对话,对应1300小时音频和参考病历。评估当前开源模型发现,级联式方法仍显著优于端到端模型。

原文摘要 · Abstract (English)

Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.

语音生成医疗AI合成数据长序列理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。