用可学习的逻辑规则实现可解释的患者活动识别,让模型会推理为何危险。
Logi-PAR: Logic-Infused Patient Activity Recognition via Differentiable Rule
- 融合多视角视觉线索,通过可微逻辑规则显式建模活动因果
- 在VAST和OmniFall数据集上超越主流视觉语言模型,风险判断准确率提升显著
- 生成可审计的推理路径,支持反事实推演(如辅助使风险降65%)
临床环境中的患者活动识别(PAR)利用活动数据提升安全与照护质量。现有模型多仅识别当前活动,依赖全局与局部注意力组合稀疏视觉线索,但仅隐式学习逻辑模式。为提升临床安全,需能推断视觉线索为何构成风险,并通过显式逻辑进行组合推理。为此,我们提出首个逻辑注入型患者活动识别框架Logi-PAR,将上下文事实融合作为多视角特征提取器,并注入神经引导的可微逻辑规则。该方法自动从视觉线索中学习规则,端到端优化,同时在训练中显式标注隐含模式。据我们所知,Logi-PAR是首个通过可学习逻辑规则对符号映射进行患者活动识别的框架。它可生成可审计的推理路径(规则轨迹),并支持反事实干预(如若提供协助,风险可降低65%)。在临床基准数据集VAST和OmniFall上的广泛评估表明,其性能达到最先进水平,显著优于视觉-语言模型与Transformer基线。
原文摘要 · Abstract (English)
Patient Activity Recognition (PAR) in clinical settings uses activity data to improve safety and quality of care. Although significant progress has been made, current models mainly identify which activity is occurring. They often spatially compose sub-sparse visual cues using global and local attention mechanisms, yet only learn logically implicit patterns due to their neural-pipeline. Advancing clinical safety requires methods that can infer why a set of visual cues implies a risk, and how these can be compositionally reasoned through explicit logic beyond mere classification. To address this, we proposed Logi-PAR, the first Logic-Infused Patient Activity Recognition Framework that integrates contextual fact fusion as a multi-view primitive extractor and injects neural-guided differentiable rules. Our method automatically learns rules from visual cues, optimizing them end-to-end while enabling the implicit emergence patterns to be explicitly labelled during training. To the best of our knowledge, Logi-PAR is the first framework to recognize patient activity by applying learnable logic rules to symbolic mappings. It produces auditable why explanations as rule traces and supports counterfactual interventions (e.g., risk would decrease by 65% if assistance were present). Extensive evaluation on clinical benchmarks (VAST and OmniFall) demonstrates state-of-the-art performance, significantly outperforming Vision-Language Models and transformer baselines. The code is available via: https://github.com/zararkhan985/Logi-PAR.git}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。