用多智能体协作让隐私政策问答更公平,支持方言用户
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems
- 设计双智能体系统:方言转换+领域知识增强
- 零样本准确率提升至0.601(原0.394),跨数据集有效
- 无需重训练,适合各类模型与场景使用
隐私政策关乎数据收集与使用,但其复杂性限制了不同群体的可及性。现有隐私政策问答(QA)系统在英语方言间表现差异明显,非标准方言使用者处于劣势。本文提出一种受以人为本设计启发的多智能体框架,以缓解方言偏差。该方法包含方言智能体,将查询转化为标准美式英语(SAE)同时保留方言语义;以及隐私政策智能体,利用领域知识优化预测结果。相比以往方法,本方案无需重新训练或方言微调,具备广泛适用性。在PrivacyQA和PolicyQA上的评估显示,GPT-4o-mini的零样本准确率从0.394提升至0.601(PrivacyQA),从0.352提升至0.464(PolicyQA),优于或匹配少样本基线,且无需额外训练数据。结果表明,结构化智能体协作能有效缓解方言偏差,强调了在NLP系统中考虑语言多样性对实现隐私信息公平获取的重要性。
原文摘要 · Abstract (English)
Privacy policies inform users about data collection and usage, yet their complexity limits accessibility for diverse populations. Existing Privacy Policy Question Answering (QA) systems exhibit performance disparities across English dialects, disadvantaging speakers of non-standard varieties. We propose a novel multi-agent framework inspired by human-centered design principles to mitigate dialectal biases. Our approach integrates a Dialect Agent, which translates queries into Standard American English (SAE) while preserving dialectal intent, and a Privacy Policy Agent, which refines predictions using domain expertise. Unlike prior approaches, our method does not require retraining or dialect-specific fine-tuning, making it broadly applicable across models and domains. Evaluated on PrivacyQA and PolicyQA, our framework improves GPT-4o-mini's zero-shot accuracy from 0.394 to 0.601 on PrivacyQA and from 0.352 to 0.464 on PolicyQA, surpassing or matching few-shot baselines without additional training data. These results highlight the effectiveness of structured agent collaboration in mitigating dialect biases and underscore the importance of designing NLP systems that account for linguistic diversity to ensure equitable access to privacy information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。