研究大模型对生理信号衍生数据的过度信任问题,提出量化与缓解方案。
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence

- 构建衍生特征过信任评估框架,通过心电图监督心率信号可信度。
- 实验显示改进方法在修复率和特异性上提升1.82至6.69个百分点。
- 适合关注大模型可靠性、医疗智能诊断的开发者与研究者。
衍生测量在大型语言模型(LLM)流程中常被当作直接事实使用,但其有效性具有实例依赖性。本文定义了衍生特征过信任(DFOT),即下游模型将此类测量误认为直接事实或超出其有效范围使用。以生理传感为例,实验D1测试模型对与离线心电图矛盾的光电容积脉搏波(PPG)心率的接受程度;实验D2测试对离线验证可靠的PPG心率在误导性病史下被拒绝的情况。心电图提供训练监督和离线参考,但从不输入给模型。五个指标量化该链条:冲突过信任率(COTR)和情境诱发错误率(CIR)分别刻画D1/D2;正确修复率(CRR)衡量错误冻结修复能力;证据特定修复边际(ESRM)对比匹配与患者无关的随机证据;效用损害率(UHR)衡量高可信度情况下本无需验证却额外验证的比例。框架不依赖特定可靠性生成器。在5万对PPG-ECG记录上,采用心电图到PPG的特权蒸馏作为基线,实现五项修复与特异性指标提升1.82至6.69个百分点,所有配对置信区间不含零;UHR上升0.67个百分点(95%置信区间:-0.4至+1.7)。该框架为更强的缓解方法提供统一评估目标。代码已开源。
原文摘要 · Abstract (English)
Derived measurements increasingly enter large language model (LLM) pipelines as direct facts despite their instance-dependent validity. We define derived-feature over-trust (DFOT) as the failure in which a downstream LLM assigns such a measurement the epistemic status of a direct fact or uses it outside its valid scope. Using physiological sensing as a case study, D1 tests acceptance of a PPG-derived rhythm contradicted by offline ECG, whereas D2 tests rejection of an offline-confirmed reliable PPG rhythm under misleading severe history. ECG supplies training supervision and offline reference construction but is never shown to the LLM. Five estimands quantify this chain: conflict over-trust rate (COTR) and context-induced error rate (CIR) characterize D1/D2; correct repair rate (CRR) measures frozen-error repair; evidence-specific repair margin (ESRM) contrasts matched and patient-disjoint shuffled evidence; and utility harm rate (UHR) measures unnecessary verification among HIGH-reliability cases used without verification at baseline. The framework does not depend on a particular reliability generator. We demonstrate it on 50,000 paired PPG-ECG records using ECG-to-PPG privileged distillation as an illustrative baseline and PPG-only inference. On a protocol-locked 187-patient test, the baseline improves four repair and specificity endpoints by 1.82-6.69 percentage points, with all paired confidence intervals excluding zero; UHR increases by 0.67 percentage points (95% CI: -0.4 to +1.7). DFOT provides a common evaluation target for stronger mitigation methods. The code is available at https://github.com/Zongheng-Guo/When-Derived-Measurements-Mislead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。