评测大模型的准确、无毒和公平性,发现其常见错误与偏见。
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
- 设计测试框架评估事实正确性与逻辑推理能力
- 通过红队攻击检测模型生成内容的毒性风险
- 提出新方法衡量社会与文化偏见,适合安全研究者
近年来,以ChatGPT为代表的大语言模型因其卓越的对话能力和智能表现,迅速渗透至工作与日常生活中。ChatGPT成为人类历史上用户增长最快的软件,并成为下一代人工智能应用的重要基础模型。然而,大模型生成的内容常存在事实错误、偏见和毒性问题。鉴于其庞大的用户基数和广泛的应用场景,这些不可靠输出可能带来严重负面影响。本论文在博士研究期间探索了大语言模型可靠性问题,从软件测试与自然语言处理双重视角,聚焦于模型的正确性、无毒性和公平性。首先,为评估事实正确性,提出FactChecker与LogicAsker两个测试框架,分别用于检验知识准确性与逻辑推理能力。其次,针对非毒性问题,引入两种红队攻击方法对模型进行对抗性测试。最后,为评估公平性,提出BiasAsker与XCulturalBench两个评估框架,分别衡量社会偏见与文化偏见。研究揭示了当前主流大模型在多维度上的系统性缺陷。
原文摘要 · Abstract (English)
Large language models (LLMs), such as ChatGPT, have rapidly penetrated into people's work and daily lives over the past few years, due to their extraordinary conversational skills and intelligence. ChatGPT has become the fastest-growing software in terms of user numbers in human history and become an important foundational model for the next generation of artificial intelligence applications. However, the generations of LLMs are not entirely reliable, often producing content with factual errors, biases, and toxicity. Given their vast number of users and wide range of application scenarios, these unreliable responses can lead to many serious negative impacts. This thesis introduces the exploratory works in the field of language model reliability during the PhD study, focusing on the correctness, non-toxicity, and fairness of LLMs from both software testing and natural language processing perspectives. First, to measure the correctness of LLMs, we introduce two testing frameworks, FactChecker and LogicAsker, to evaluate factual knowledge and logical reasoning accuracy, respectively. Second, for the non-toxicity of LLMs, we introduce two works for red-teaming LLMs. Third, to evaluate the fairness of LLMs, we introduce two evaluation frameworks, BiasAsker and XCulturalBench, to measure the social bias and cultural bias of LLMs, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。