arXiv:2511.10297cs.CL2025-11被引 1

本地运行的智能文档问答系统,兼顾隐私与高准确率。

Local Hybrid Retrieval-Augmented Document QA

  • 融合语义与关键词检索,本地化部署无须联网。
  • 在法律、科学和对话类文档上实现高精度问答。
  • 适合银行、医院、律所等对数据隐私要求高的机构。

处理敏感文档的组织面临两难:采用云端AI虽有强大问答能力却牺牲数据隐私,本地部署则准确性差。我们提出一种完全本地运行的问答系统,结合语义理解与关键词精准匹配,在无需互联网连接的本地基础设施上运行。该方法在法律、科学和对话类文档上实现了与云端系统相媲美的复杂查询准确率,同时确保所有数据保留在企业内部。通过平衡两种互补的检索策略并利用消费级硬件加速,系统以极低错误率提供可靠答案,使金融机构、医疗机构和律师事务所可在不外传机密信息的前提下使用对话式文档AI。本研究证明,企业级AI部署中隐私与性能不必相互妥协。

原文摘要 · Abstract (English)

Organizations handling sensitive documents face a critical dilemma: adopt cloud-based AI systems that offer powerful question-answering capabilities but compromise data privacy, or maintain local processing that ensures security but delivers poor accuracy. We present a question-answering system that resolves this trade-off by combining semantic understanding with keyword precision, operating entirely on local infrastructure without internet access. Our approach demonstrates that organizations can achieve competitive accuracy on complex queries across legal, scientific, and conversational documents while keeping all data on their machines. By balancing two complementary retrieval strategies and using consumer-grade hardware acceleration, the system delivers reliable answers with minimal errors, letting banks, hospitals, and law firms adopt conversational document AI without transmitting proprietary information to external providers. This work establishes that privacy and performance need not be mutually exclusive in enterprise AI deployment.

文档问答本地推理隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。