arXiv:2607.27245cs.SDcs.AI2026-07

用轻量微调让Whisper更准转录执法录像,还能在普通电脑上跑。

Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage

论文配图:Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
图 1 · 摘自论文原文
  • 用LoRA微调Whisper,仅改少量参数就适配执法场景音频。
  • 在4GB显卡上用8比特量化实现高效推理,准确率达93.7%。
  • 适合需要低成本自动化转录的警务系统与司法透明研究者。

现代执法面临“可见性悖论”:警员佩戴摄像头(BWC)积累了海量视频资料,却因人工转录成本过高而难以用于问责或系统性审查。本研究提出一种框架,将OpenAI Whisper模型适配至执法环境特有的语音挑战。通过低秩适应(LoRA)实现参数高效微调,有效缓解零样本模型在高压情境、警笛和无线电干扰下的性能下降问题。关键成果在于,该方法可在消费级硬件(搭载NVIDIA 4GB GTX GPU的Acer Nitro笔记本)上完成训练与推理,结合8位量化与梯度检查点技术。进一步将转录结果接入领域特定本体的符号推理流程,生成证据关联的事件图谱,实现93.7%的词汇映射率,助力程序正义与透明化治理。

原文摘要 · Abstract (English)

Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI Whisper architecture to the unique acoustic and linguistic challenges of the policing environment. By employing Parameter-Efficient Fine-Tuning (PEFT) through Low-Rank Adaptation (LoRA), we address the significant performance degradation observed in zero-shot models when confronted with high-stress scenarios, sirens, and radio interference. Crucially, we demonstrate that this adaptation is feasible on consumer-grade hardware (Acer Nitro local machine with NVIDIA 4GB GTX GPU) using 8-bit quantization and gradient checkpointing. We further integrate these transcriptions into a symbolic reasoning pipeline using a domain-specific ontology to transform raw audio into evidence-linked incident graphs, achieving a 93.7% lexicon mapping rate for the advancement of procedural justice and transparency.

语音识别LoRA执法科技Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。