构建可审计的电商行为推断框架,用结构化元素提升透明度与可复现性。
From Explicit Elements to Implicit Intent: A Predefined Library for Auditable Behavioral Inference
- 基于24个行为元素分层架构,通过元数据驱动可插拔推断
- 实现完全可复现(sigma=0)和可控变异性推断,精度损失换取透明性
- 适合需要可解释决策链的合规性或风控场景
我们提出SemantiClean,一个模块化框架,从电商平台会话数据中提取结构化语义信号,通过共享元素库支持购买意图、客户分群和产品偏好等可插拔推断目标。与仅追求准确率的传统端到端模型不同,SemantiClean优先保障可审计性、结构治理和sigma=0的可复现性,明确以小幅预测性能为代价换取元素级透明与可辩护的决策路径。基于Online Shoppers Purchasing Intention (OSPI)数据集,该框架将24个行为元素组织为四层架构(功能层、交互层、系统层、上下文层),并通过三项反通胀机制确保信号质量:冗余组贡献上限、分级惩罚计算器偏差惩罚、自适应约束模式冷启动保护。本报告引入了集成大模型的语义推断引擎,采用两阶段大模型驱动架构,在推理时利用完整的元素元数据。所有定量结果均由该引擎生成。确定性输出保持完全可复现(sigma=0);大模型依赖结果(E8, E10)在固定提供方/模型/温度设置下具有受控变异性。当前实现中性别推断目标仍不可用,未参与任何定量评估。
原文摘要 · Abstract (English)
We present SemantiClean, a modular framework for extracting structured semantic signals from e-commerce session data and driving pluggable inference targets including purchase intent, customer segmentation, and product affinity through a shared element library. Unlike conventional end-to-end predictors that optimise solely for accuracy, SemantiClean prioritises auditability, structural governance, and sigma=0 reproducibility, explicitly trading marginal predictive gains for element-level transparency and defensible decision trails. Built upon the Online Shoppers Purchasing Intention (OSPI) dataset, the framework organises twenty-four behavioural elements into a four-layer architecture (Functional, Interaction, Systemic, Contextual) and enforces signal quality through three anti-inflation mechanisms: RedundancyGroup contribution caps, TieredPenaltyCalculator bias penalties, and AdaptiveConstraintMode cold-start protection.This report introduces the LLM-Integrated Semantic Inference Engine, a fully implemented two-phase LLM-driven inference architecture that leverages complete element metadata at inference time. All quantitative results reported herein are produced by this engine. Deterministic engine outputs remain fully reproducible (sigma=0); LLM-dependent results (E8, E10) are subject to controlled output variability under fixed provider/model/temperature settings. The gender inference target remains non-functional in the current implementation and is excluded from all quantitative results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。