打造可追溯的浏览器语音采集系统,确保远程语音研究数据可信
Beyond .WAV: Design and Software Verification of VocalCap, a Traceable Browser-Based Audio Capture System for Vocal Biomarker Research
- 用版本化协议驱动浏览器端语音采集,全程记录技术细节
- 14个左声道文件经主动通道选择,保真度误差小于0.001 dB
- 适合临床语音生物标志物研究,需配合设备与人群验证
远程语音研究常仅保留音频文件,缺乏采集、传输、处理和接受的完整证据。本文提出VocalCap,一种机构可控、无需技术培训的浏览器端语音及声学信号自引导采集系统。通过版本化协议驱动流程,每个接受的录音均保留浏览器原生对象、基于同一MediaStream生成的无损Float32 WAV,以及服务器标准的单通道PCM16 WAV,关联采集执行、技术质量、字节级完整性、恢复与转换溯源证据。IndexedDB在服务器确认前保存已接受的浏览器产物,会话完成需通过所有任务与产物的验证。软件测试通过伪造或篡改对象、零值中断、通道拓扑变异、中断/重复操作等挑战采集契约。对39份同意试点录音的后验技术审计发现:25个样本完全相同的立体声文件,14个信号仅限左声道。采用拓扑感知的主动通道选择,使14个异常文件的校准均方根电平差低于0.001 dB;若采用等权立体声平均,将引入约6.02 dB衰减。生产环境端到端验证在Chromium与WebKit中完成两个五任务流程,生成10个接受录音与30个保留产物,全部通过服务器侧完整性与格式检查。结果验证了VocalCap在所测浏览器引擎条件下的软件行为。设备级声学一致性、目标人群可用性、临床有效性及生物标志物性能仍需独立研究。
原文摘要 · Abstract (English)
Remote voice studies often retain a final audio file with limited evidence about how it was captured, transferred, processed, and accepted. This paper presents VocalCap, an institution-controlled, browser-based system for self-guided capture of voice and related acoustic signals by participants without technical training. A versioned protocol drives the workflow. Each accepted recording retains a browser-native object, a client-lossless Float32 WAV derived from the same MediaStream, and a server-canonical mono PCM16 WAV, linked to evidence of capture execution, technical quality, byte-level integrity, recovery, and transformation provenance. IndexedDB preserves accepted browser artifacts until server confirmation, while session completion requires successful verification of every task and artifact. Software tests challenged the acquisition contracts with malformed or altered objects, exact-zero interruptions, channel-topology variants, and interrupted or repeated operations. A post hoc technical audit of 39 consented pilot recordings found 25 sample-identical stereo files and 14 files with signal confined to the left channel. Topology-aware active-channel selection limited the canonical root-mean-square level difference to less than 0.001 dB in all 14 affected files; equal-weight stereo averaging would have introduced approximately 6.02 dB of attenuation. Production end-to-end verification completed two five-task profiles in Chromium and WebKit, yielding 10 accepted recordings and 30 retained artifacts that passed server-side integrity and format checks. The results verify VocalCap's software behavior under the tested browser-engine conditions. Device-level acoustic agreement, target-population usability, clinical validity, and biomarker performance remain subjects for separate studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。