打造合规房地产聊天机器人,避免歧视性推荐。
A Recipe For Building a Compliant Real Estate Chatbot
- 用合成数据生成指令遵循与安全行为训练集。
- 微调后模型性能媲美GPT-4o,且更安全合规。
- 适合关注公平性与行业应用的AI研究者。
近年来,大型语言模型与人类偏好对齐取得显著进展。本文聚焦于构建一个面向房地产领域的专用聊天机器人,重点引入合规行为以防止歧视性做法(如引导和红线政策),这些现象曾在美国房地产行业长期存在。基于已有工作,我们提出一种生成合成通用指令遵循数据集及安全数据的方法。通过大量评估与基准测试,我们对 llama-3-8B-instruct 模型进行微调,显著提升其性能,使其在合规性与安全性上达到与 GPT-4o 等闭源大模型相当的水平。相关模型、数据与代码已开源,以支持社区进一步研究与发展。
原文摘要 · Abstract (English)
In recent years, there has been significant effort to align large language models with human preferences. This work focuses on developing a chatbot specialized in the real estate domain, with an emphasis on incorporating compliant behavior to ensure it can be used without perpetuating discriminatory practices like steering and redlining, which have historically plagued the real estate industry in the United States. Building on prior work, we present a method for generating a synthetic general instruction-following dataset, along with safety data. Through extensive evaluations and benchmarks, we fine-tuned a llama-3-8B-instruct model and demonstrated that we can enhance it's performance significantly to match huge closed-source models like GPT-4o while making it safer and more compliant. We open-source the model, data and code to support further development and research in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。