构建首个公开的房产公平性数据集,用于检测大模型在地产交易中的合规风险。
FairHome: A Fair Housing and Fair Lending Dataset
- 基于7.5万条标注数据,构建房产领域合规风险二分类数据集。
- 训练分类器在零样本和少样本下对齐大模型行为,F1达0.91。
- 适合监管科技、算法审计与公平性研究者使用。
我们提出一个名为FairHome的公平住房与公平贷款数据集:包含约7.5万条样本,覆盖9个受保护类别。据我们所知,FairHome是首个在住房领域提供合规风险二分类标签的公开数据集。通过训练分类器,并用于检测大语言模型(LLM)在房地产交易场景下的潜在违规行为,验证了该数据集的有效性。我们在零样本和少样本条件下,将训练好的分类器与GPT-3.5、GPT-4、LLaMA-3和Mistral Large等先进大模型进行对比。结果表明,分类器F1分数达到0.91,充分证明了该数据集的价值。
原文摘要 · Abstract (English)
We present a Fair Housing and Fair Lending dataset (FairHome): A dataset with around 75,000 examples across 9 protected categories. To the best of our knowledge, FairHome is the first publicly available dataset labeled with binary labels for compliance risk in the housing domain. We demonstrate the usefulness and effectiveness of such a dataset by training a classifier and using it to detect potential violations when using a large language model (LLM) in the context of real-estate transactions. We benchmark the trained classifier against state-of-the-art LLMs including GPT-3.5, GPT-4, LLaMA-3, and Mistral Large in both zero-shot and few-shot contexts. Our classifier outperformed with an F1-score of 0.91, underscoring the effectiveness of our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。