PolicyBot让普通人也能准确读懂政策文件中的关键信息。
PolicyBot - Reliable Question Answering over Policy Documents
- 用领域特化的分块与多语言向量检索,精准定位政策内容
- 通过溯源机制减少幻觉,确保回答可验证、可追溯
- 全开源设计,适合政府或组织部署到其他法规问答场景
一个国家的公民受其政府所制定法律与政策的影响。这些法规对公民具有重要意义,如赋予权利或设定义务。然而,政策文件通常冗长复杂、难以查找和理解,导致公众难以获取所需信息。本文提出PolicyBot,一种聚焦透明性与可复现性的检索增强生成(RAG)系统,用于在政策文档上回答用户问题。该系统结合领域特定的语义分块、多语言稠密嵌入、多阶段检索与重排序,以及源感知生成,确保回答基于原始文档。我们引入引用追踪机制以减少幻觉,提升用户信任,并评估了不同检索与生成配置,识别出有效设计选择。整个端到端流程完全使用开源工具构建,便于扩展至其他需要文档驱动问答的领域。本工作强调了在治理类场景中部署可信RAG系统的设计考量、实际挑战与经验教训。
原文摘要 · Abstract (English)
All citizens of a country are affected by the laws and policies introduced by their government. These laws and policies serve essential functions for citizens. Such as granting them certain rights or imposing specific obligations. However, these documents are often lengthy, complex, and difficult to navigate, making it challenging for citizens to locate and understand relevant information. This work presents PolicyBot, a retrieval-augmented generation (RAG) system designed to answer user queries over policy documents with a focus on transparency and reproducibility. The system combines domain-specific semantic chunking, multilingual dense embeddings, multi-stage retrieval with reranking, and source-aware generation to provide responses grounded in the original documents. We implemented citation tracing to reduce hallucinations and improve user trust, and evaluated alternative retrieval and generation configurations to identify effective design choices. The end-to-end pipeline is built entirely with open-source tools, enabling easy adaptation to other domains requiring document-grounded question answering. This work highlights design considerations, practical challenges, and lessons learned in deploying trustworthy RAG systems for governance-related contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。