让人类参与NLP模型安全评估,提升可信度与可靠性。
From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

- 引入人机协同方法,增强模型审计与验证能力
- 发现低资源语言和私有数据库场景下评估难、成本高
- 适合关注AI安全、可解释性与合规部署的研究者
大型语言模型广泛应用于高风险自然语言处理任务,但仍存在偏见、幻觉、对抗脆弱性和不可靠泛化等风险。基于探测的审计揭示了模型行为的一致性问题,对抗文本生成暴露了鲁棒性缺陷,尤其在资源有限的语言中缺乏基准。企业级文本到SQL任务凸显了对私有大规模数据库输出验证的困难。人类监督对于探测验证、对抗检验和领域特定标注至关重要,但成本高且难以扩展。本综述探讨了近期人机协同方法,推动NLP从自动化向协作式安全与可信发展。我们梳理了人类专家如何支持审计、鲁棒性评估、数据构建与模型引导。研究发现,在可扩展探测、可持续鲁棒性基准、低资源环境及私有系统治理方面仍存空白。本文提出自适应审计、协作评估与可问责部署等实践研究方向。
原文摘要 · Abstract (English)
Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable generalization remain. Probe-based auditing reveals inconsistencies in model behavior. Adversarial text generation uncovers robustness gaps, especially in lower-resourced languages with limited benchmarks. Enterprise text-to-SQL settings expose the difficulty of validating outputs over private and large-scale databases. Human supervision is essential for probe validation, adversarial verification and domain-specific annotation, but it is costly and hard to scale. This survey examines recent human-in-the-loop methods that shift NLP from automation toward collaboration for safety and trustworthiness. We review how human expertise supports auditing, robustness evaluation, data construction and model steering. Our findings highlight gaps in scalable probing, sustainable robustness benchmarks, low-resource settings and governance of private systems. We outline practical research directions for adaptive auditing, collaborative evaluation and accountable deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。