arXiv:2508.15801cs.CLcs.AI2025-08Conference of the …

构建语音通话实体识别基准,提升大模型在真实对话中的鲁棒性。

LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts

  • 用大模型生成带语言变异的合成数据,覆盖停顿、重叠等复杂场景
  • 优化提示后在真实通话上达到94-95%的实体识别准确率
  • 适合需低成本部署且重视口语变体适应性的医疗与客服场景

我们研究客户支持和医疗场景中电话通话转录文本的结构化实体抽取问题,由于标注成本高且受隐私和同意限制,真实数据难以共享。现有方法在存在不流畅、打断和说话人重叠时性能下降。为此,我们提出LingVarBench,一个基准和语义合成数据生成流程,通过(1)大模型采样的实体值,(2)涵盖多样不流畅和特定实体读法的精心设计语言表达模式,(3)值-转录一致性过滤器生成语言多样化的训练数据。利用该数据集,DSPy的SIMBA自动合成并优化抽取提示,减少人工提示工程,增强对语言变化的鲁棒性。在真实客户通话数据上,仅基于LingVarBench优化的提示优于零样本基线,对邮政编码、出生日期、姓名等结构化实体的F1达到约94-95%;对主观问卷项,优化提示显著超越零样本表现,接近人工调优水平。LingVarBench为直接答案场景提供了实用且低成本的部署路径,后续真实标注可进一步优化模型。

原文摘要 · Abstract (English)

We study structured entity extraction from phone-call transcripts in customer-support and healthcare settings, where annotation is costly, and data access is limited by privacy and consent. Existing methods degrade under disfluencies, interruptions, and speaker overlap, yet large real-call corpora are rarely shareable. We introduce LingVarBench, a benchmark and semantic synthetic data generation pipeline that generates linguistically varied training data via (1) LLM-sampled entity values, (2) curated linguistic verbalization patterns covering diverse disfluencies and entity-specific readout styles, and (3) a value-transcript consistency filter. Using this dataset, DSPy's SIMBA automatically synthesizes and optimizes extraction prompts, reducing manual prompt engineering and targeting robustness to verbal variation. On real customer transcripts, prompts optimized solely on LingVarBench outperform zero-shot baselines and match or closely approach human-tuned prompts for structured entities such as ZIP code, date of birth, and name (F1 approximately 94-95 percent). For subjective questionnaire items, optimized prompts substantially improve over zero-shot performance and approach human-tuned prompts. LingVarBench offers a practical and cost-efficient path to deployment in a direct-answer setting, with real annotations later enabling additional refinement.

实体识别语音转录大模型应用合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。