LLM在智能系统中常偏信用户言语而忽视传感器数据,导致决策失效。
Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors
- 发现语言模型在融合数据时存在格式依赖,数值传感器数据难被采纳
- 提出两种可计算审计指标,验证了传感器信任度接近零(AAI=-0.805)
- 设计实时干预方法GAC,显著提升人体行为识别准确率至21.9%~27.5%
大型语言模型(LLMs)越来越多地融合异构输入用于泛在系统。然而,当传感器测量值与用户陈述冲突时,模型如何隐式分配权威性仍缺乏研究,这在物理感知需优先的部署场景中引发严重可靠性问题。不同于显式的传统融合,LLM将权威分配隐藏于学习表征中。我们发现这种分配严重依赖数据格式:数值型传感器数据难以融入与答案相关模型方向,导致自然语言声明主导最终决策,这一现象称为**权威倒置**。为诊断和缓解此问题,我们提出一个几何上下文融合框架,引入两个可计算的审计指标——上下文融合比(CIR)与权威对齐指数(AAI),并提出一种推理时层级干预方法——几何权威校准(GAC),以抑制错误的用户权威。在四个数据集共576个冲突实例上评估四个模型(4B至35B参数,三种架构),结果显示极端倒置:在数值任务中,模型对传感器的信任近乎为零(AAI = -0.805,Cohen's d = -2.14),且不受模型容量影响。验证几何框架后,理论指导的因果注入使80.2%的错误决策得以纠正(随机控制<0.4%)。实践中,GAC将人体活动识别(HAR)准确率从0–1.6%提升至21.9–27.5%,优于提示基线。最终表明,LLM系统中的权威分配必须显式审计并按应用配置,而非默认隐含。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly fuse heterogeneous inputs in ubiquitous systems. Yet, how LLMs implicitly allocate authority when sensor measurements and user claims conflict remains unexamined, raising critical reliability concerns for deployments where physical sensing must retain priority. Unlike explicit traditional fusion, LLMs bury authority allocation within learned representations. We discover this allocation is severely format-dependent: numerical sensor data fails to integrate into answer-relevant model directions, allowing natural-language claims to dominate the final decision, a phenomenon we term \textbf{Authority Inversion}.To diagnose and mitigate this, we develop a geometric framework of context integration, introduce two computable audit metrics, specifically the Context Integration Ratio (CIR) and Authority Alignment Index (AAI), and propose Geometric Authority Calibration (GAC), an inference-time layer-level intervention to suppress misplaced user authority. Evaluating four models (4B to 35B parameters, three architectures) across four datasets totaling 576 conflict instances reveals extreme inversion: on numerical tasks, models exhibit near-zero sensor trust (AAI = -0.805, Cohen's d = -2.14), unaffected by model capacity. Validating our geometric framework, theory-guided causal injection flips 80.2\% of incorrect decisions (vs. <0.4\% for random controls). Practically, GAC improves HAR accuracy from 0 -- 1.6\% to 21.9 -- 27.5\%, outperforming prompting baselines. Ultimately, authority allocation in LLM-mediated systems must be explicitly audited and application-specifically configured rather than left implicit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。