arXiv:2502.06652cs.CL2025-02被引 2

用RAG与对齐技术提升隐私问答的透明度和可理解性。

Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A

  • 结合RAG与对齐模块,增强回答的准确性和可读性。
  • 多维度扩展的MultiRAIN在21项指标中优于基线系统。
  • 适合法律合规、隐私保护等需要透明AI的应用场景。

GDPR的透明性原则要求数据处理信息清晰、精确且可访问。尽管语言模型在此领域具有潜力,但其概率特性使真实性与可理解性面临挑战。本文研究了结合对齐技术的前沿检索增强生成(RAG)系统,以满足GDPR要求。我们评估了集成对齐模块(如可回溯自回归推理RAIN)及提出的多维扩展MultiRAIN的RAG系统,使用隐私问答数据集进行测试。响应经过精确性与可理解性优化,并通过21项指标评估,包括确定性和大模型评测。结果显示,带对齐模块的RAG系统在多数指标上优于基线,但尚未完全达到人类回答水平。主成分分析揭示各指标间存在复杂交互,提示需进一步优化评估体系。本研究为将先进自然语言处理系统融入法律合规框架奠定基础。

原文摘要 · Abstract (English)

The transparency principle of the General Data Protection Regulation (GDPR) requires data processing information to be clear, precise, and accessible. While language models show promise in this context, their probabilistic nature complicates truthfulness and comprehensibility. This paper examines state-of-the-art Retrieval Augmented Generation (RAG) systems enhanced with alignment techniques to fulfill GDPR obligations. We evaluate RAG systems incorporating an alignment module like Rewindable Auto-regressive Inference (RAIN) and our proposed multidimensional extension, MultiRAIN, using a Privacy Q&A dataset. Responses are optimized for preciseness and comprehensibility and are assessed through 21 metrics, including deterministic and large language model-based evaluations. Our results show that RAG systems with an alignment module outperform baseline RAG systems on most metrics, though none fully match human answers. Principal component analysis of the results reveals complex interactions between metrics, highlighting the need to refine metrics. This study provides a foundation for integrating advanced natural language processing systems into legal compliance frameworks.

隐私问答RAG对齐技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。