在真实语料上测试提示词上下文对语音转录的影响,发现无显著提升。
No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

- 在生产级语音转录工具中测试完整提示上下文,采用配对消融设计。
- 全上下文与无上下文相比,侧级别词错误率(WER)差异不显著,平均+0.6点。
- 需结合逐词、插入词和说话人标签等细粒度指标评估上下文效果。
在生产级口述历史转录工具的提示条件层中,测试了推理时提供上下文这一低成本适配机制的效果。研究基于其自有生产语料的样本,采用预注册的配对消融实验设计,分析代码在评分前已通过哈希冻结。共处理19段录音带(约10.6小时),覆盖1970–1980年代退化音频,使用gpt-4o-transcribe与gemini-2.5-flash两种部署配置,在三种提示策略下重新处理,并与人工校正的逐字参考文本比对。对gpt-4o-transcribe,全上下文与无上下文的中位差为+0.6 WER点,侧级重抽样区间为[-1.1, +1.0];Gemini结果因不稳定性无法得出可靠结论。事后分析显示,运行间管道变异大于确认差异,因此单次转录无法分辨此类效应。实施审计确认操作生效,序列对齐分析发现仅完整列出短语有微小改善,不足以影响侧级整体WER,而Gemini还伴随未列出词错误增加。因此,评估上下文机制需结合逐词、插入词及说话人标签等序列对齐指标。
原文摘要 · Abstract (English)
Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the prompt-conditioning layer of a production oral-history transcription tool, on a sample from its own production corpus. Full prompt-level context did not detectably change side-level word error rate (WER), and none of the four preregistered hypotheses was supported. The design was a within-item paired ablation, preregistered with the analysis code frozen by hash before the confirmatory batch was scored; two disclosed gpt-4o pilot sides had been scored earlier, during scorer development. Nineteen cassette sides, about 10.6 hours of degraded 1970s-80s interview audio, were reprocessed through the production code path under three prompt arms, crossed with two deployed commercial configurations, gpt-4o-transcribe and gemini-2.5-flash, and scored against operator-corrected verbatim references. For gpt-4o-transcribe the median paired difference between the full-context and no-context arms was +0.6 WER points, with a side-resampled interval of [-1.1, +1.0]; the Gemini estimates were too unstable to support a comparable negative inference. A post-hoc rerun found run-to-run pipeline variability larger than the confirmatory differences, so effects of that size cannot be resolved from one transcription per cell. An implementation audit verified the manipulation was live, and sequence-alignment analysis found a small improvement on complete context-listed phrases, too small to materially change side-level WER, and for Gemini coexisting with worsened unlisted-token error. Evaluating context mechanisms therefore requires sequence-aligned term-level, insertion, and speaker-label measures alongside aggregate accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。