arXiv:2604.24700cs.CLcs.AI2026-04

研究用户提问微小变化如何影响大模型输出,提出可信AI部署新方法。

Green Shielding: A User-Centric Approach Towards Trustworthy AI

论文配图:Green Shielding: A User-Centric Approach Towards Trustworthy AI
图 1 · 摘自论文原文
  • 以真实患者提问构建医疗诊断基准,评估模型行为变化。
  • 发现提示词调整会系统改变诊断结果的全面性与临床合理性。
  • 适合医疗、高风险决策等领域的AI安全部署参考。

大型语言模型(LLMs)应用日益广泛,但其输出对用户提问方式的微小非对抗性变化高度敏感,现有红队测试未能有效解决此问题。本文提出「Green Shielding」——一种以用户为中心的可信AI构建方法,通过分析良性输入变化如何影响模型行为,提供实证支持的部署指南。该方法基于CUE标准:包含真实上下文、参考标准与衡量实际效用的指标,以及反映真实提问差异的扰动。结合PCS框架并由执业医生参与,我们在医疗诊断场景中构建了HealthCareMagic-Diagnosis(HCM-Dx)基准,包含患者原创提问、结构化参考诊断集及临床导向评估指标。研究不同扰动策略发现,常规提问变化可使模型行为沿临床相关维度发生系统性偏移。在多个前沿大模型中,中性化处理(去除常见用户特征但保留临床内容)提升诊断合理性,生成更简洁、类医者风格的鉴别诊断列表,但降低了对高概率及关键安全病症的覆盖。结果表明,交互选择能系统影响输出任务相关属性,为高风险领域安全部署提供用户导向指导。该方法可扩展至其他决策支持场景与代理型AI系统。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well addressed by existing red-teaming efforts. We propose Green Shielding, a user-centric agenda for building evidence-backed deployment guidance by characterizing how benign input variation shifts model behavior. We operationalize this agenda through the CUE criteria: benchmarks with authentic Context, reference standards and metrics that capture true Utility, and perturbations that reflect realistic variations in the Elicitation of model behavior. Guided by the PCS framework and developed with practicing physicians, we instantiate Green Shielding in medical diagnosis through HealthCareMagic-Diagnosis (HCM-Dx), a benchmark of patient-authored queries, together with structured reference diagnosis sets and clinically grounded metrics for evaluating differential diagnosis lists. We also study perturbation regimes that capture routine input variation and show that prompt-level factors shift model behavior along clinically meaningful dimensions. Across multiple frontier LLMs, these shifts trace out Pareto-like tradeoffs. In particular, neutralization, which removes common user-level factors while preserving clinical content, increases plausibility and yields more concise, clinician-like differentials, but reduces coverage of highly likely and safety-critical conditions. Together, these results show that interaction choices can systematically shift task-relevant properties of model outputs and support user-facing guidance for safer deployment in high-stakes domains. Although instantiated here in medical diagnosis, the agenda extends naturally to other decision-support settings and agentic AI systems.

可信AI医疗诊断大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。