小模型本地部署,帮记者安全高效查文档
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
- 用五阶段流程+小模型实现可审计的文档搜索
- 三款量化模型在桌面端均运行良好,引用准确率高
- 适合注重隐私与可控性的新闻机构使用
调查记者常需处理大量文档。尽管具备检索增强生成能力的大语言模型(LLM)有望加速文档发现,但因幻觉风险、验证负担和数据隐私问题,新闻室采用仍受限。本文提出一种以记者为中心的LLM驱动文档搜索方法,通过五阶段流程——语料摘要、搜索规划、并行线程执行、质量评估和合成——使用可本地部署的小模型,在保障数据安全的同时实现全程可审计的引用链。在两个语料库上评估了三款量化模型(Gemma 3 12B、Qwen 3 14B、GPT-OSS 20B),结果显示所有模型均在标准桌面硬件(如24 GB内存)上高效运行,引用有效性高。但系统性挑战仍存,包括多阶段合成中的错误传播,以及性能随训练数据与语料重叠度变化而显著波动。研究指出,新闻室有效部署AI需精心选型、合理设计,并辅以人工监督以确保准确性与问责性。
原文摘要 · Abstract (English)
Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate the process of document discovery, but newsroom adoption remains limited due to hallucination risks, verification burden, and data privacy concerns. We present a journalist-centered approach to LLM-powered document search that prioritizes transparency and editorial control through a five-stage pipeline -- corpus summarization, search planning, parallel thread execution, quality evaluation, and synthesis -- using small, locally-deployable language models that preserve data security and maintain complete auditability through explicit citation chains. Evaluating three quantized models (Gemma 3 12B, Qwen 3 14B, and GPT-OSS 20B) on two corpora, we find substantial variation in reliability. All models achieved high citation validity and ran effectively on standard desktop hardware (e.g., 24 GB of memory), demonstrating feasibility for resource-constrained newsrooms. However, systematic challenges emerged, including error propagation through multi-stage synthesis and dramatic performance variation based on training data overlap with corpus content. These findings suggest that effective newsroom AI deployment requires careful model selection and system design, alongside human oversight for maintaining standards of accuracy and accountability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。