arXiv:2605.14381cs.LGcs.CL2026-05

用真实社会证据生成有偏差的合成数据,让大模型暴露出更多安全漏洞。

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

论文配图:NodeSynth: Socially Aligned Synthetic Data for AI Evaluation
图 1 · 摘自论文原文
  • 基于真实证据的细粒度分类生成社会相关合成查询
  • 对主流大模型测试,失败率最高达人工基准的5倍
  • 适合评估大模型安全性和改进防护机制的研究者使用

生成式AI的发展推动了大规模合成数据用于模型评估。然而,缺乏针对性方法时,这些数据往往缺少敏感领域所需的社会技术细节。我们提出NodeSynth,一种基于真实证据的细粒度分类生成器(TaG)驱动的方法,生成具有社会相关性的合成查询。在四个主流大模型(如Claude 4.5 Haiku)上评估显示,其引发的失败率最高是人工基准的五倍。消融实验表明,细粒度分类扩展显著提升失败率;独立验证发现主流防护模型(如Llama-Guard-3)存在严重缺陷。我们开源完整研究原型与数据集,支持可扩展、高风险场景下的模型评估与定向安全干预(https://github.com/google-research/nodesynth)。

原文摘要 · Abstract (English)

Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the sociotechnical nuance required for sensitive domains. We introduce NodeSynth, an evidence-grounded methodology that generates socially relevant synthetic queries by leveraging a fine-tuned taxonomy generator (TaG) anchored in real-world evidence. Evaluated against four mainstream LLMs (e.g., Claude 4.5 Haiku), NodeSynth elicited failure rates up to five times higher than human-authored benchmarks. Ablation studies confirm that our granular taxonomic expansion significantly drives these failure rates, while independent validation reveals critical deficiencies in prominent guard models (e.g., Llama-Guard-3). We open-source our end-to-end research prototype and datasets to enable scalable, high-stakes model evaluation and targeted safety interventions (https://github.com/google-research/nodesynth).

合成数据模型评估安全检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。