构建工具链评估生成文本在医疗与法律领域的可用性与隐私风险
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
- 集成生成与多维度评估模块,支持自定义合成文本
- 覆盖下游任务性能、公平性、隐私泄露等五类评估指标
- 专为医疗与法律高风险领域设计,助力隐私保护型AI开发
我们提出 SynthTextEval,一个用于全面评估合成文本的工具包。大型语言模型输出的流畅性使合成文本在诸多场景中具备应用潜力,例如降低高风险领域中人工智能系统开发与部署的隐私泄露风险。然而,要实现这一潜力,需在多个维度上对合成数据进行一致、系统的评估:其在下游系统中的有效性、系统的公平性、隐私泄露风险、与源文本的分布差异,以及领域专家的定性反馈。SynthTextEval 支持用户上传或使用工具包内置生成模块创建的合成数据,并在所有这些维度上进行评估。尽管该工具可应用于任意数据,我们重点展示了其在医疗和法律两个高风险领域中的功能与效果。通过整合并标准化评估指标,我们旨在提升合成文本的可行性,从而促进人工智能开发中的隐私保护。
原文摘要 · Abstract (English)
We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentially viable for numerous applications, such as reducing the risks of privacy violations in the development and deployment of AI systems in high-stakes domains. Realizing this potential, however, requires principled consistent evaluations of synthetic data across multiple dimensions: its utility in downstream systems, the fairness of these systems, the risk of privacy leakage, general distributional differences from the source text, and qualitative feedback from domain experts. SynthTextEval allows users to conduct evaluations along all of these dimensions over synthetic data that they upload or generate using the toolkit's generation module. While our toolkit can be run over any data, we highlight its functionality and effectiveness over datasets from two high-stakes domains: healthcare and law. By consolidating and standardizing evaluation metrics, we aim to improve the viability of synthetic text, and in-turn, privacy-preservation in AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。