用多智能体拆解文本事实与观点,精准识别偏见并解释原因。
Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data
- 将每句话拆解为事实或观点,再评分偏见强度。
- 在WikiNPOV数据集上准确率达84.9%,比零样本提升13.0%。
- 适合关注AI公平性与可解释性的研究人员和开发者。
从AI聊天机器人传播虚假信息到推荐系统无意强化刻板印象,文本偏见严重威胁大语言模型的可信度。本文提出一种多智能体框架,通过分离每句话的事实与观点属性,赋予偏见强度评分,并提供简洁、客观的解释。在1,500个样本的WikiNPOV数据集上评估,该框架准确率达到84.9%,相比零样本基线提升13.0%,证明了在量化偏见前显式建模事实与观点的有效性。结合高精度检测与可解释性说明,为提升现代语言模型的公平性与问责制奠定基础。
原文摘要 · Abstract (English)
From disinformation spread by AI chatbots to AI recommendations that inadvertently reinforce stereotypes, textual bias poses a significant challenge to the trustworthiness of large language models (LLMs). In this paper, we propose a multi-agent framework that systematically identifies biases by disentangling each statement as fact or opinion, assigning a bias intensity score, and providing concise, factual justifications. Evaluated on 1,500 samples from the WikiNPOV dataset, the framework achieves 84.9% accuracy$\unicode{x2014}$an improvement of 13.0% over the zero-shot baseline$\unicode{x2014}$demonstrating the efficacy of explicitly modeling fact versus opinion prior to quantifying bias intensity. By combining enhanced detection accuracy with interpretable explanations, this approach sets a foundation for promoting fairness and accountability in modern language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。