arXiv:2608.11555cs.LGstat.AP2026-08中稿 · the 5th Workshop o…

提出验证信号是否提升客户退货时间预测的可靠方法。

Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough

  • 设计筛选-确认协议,通过阳性对照检验模型检测能力。
  • 发现连续衰减机制几乎可完全解释退货时间规律。
  • 适用于电商与真实场景,适合关注模型可信评估的研究者。

实践者不断为客户退货模型添加新信号(如生命周期价值、类别、最近一次行为等),时间点过程(TPP)文献也相应采用协变量和外部协变量条件强度。但这些信号是否真能改善预测?一个‘无帮助’的结论只有在模型有能力发现信号时才可信。本文提出两个贡献:(1)一种筛选-确认协议,通过预设已知强度的耦合信号验证模型能否恢复,从而确保真实数据中‘无信号’是可信的‘无效应’而非‘方法弱’;该方法在分类与连续编码下均有效,并在真实时钟驱动数据集(纽约出租车每小时数据)上验证。(2)一种无需模型的上限,衡量客户退货时间可预测性的极限——任何协变量对时间方差的解释不足一位数百分比,退货行为近乎无记忆性。利用这些工具,在三个公开基准(Amazon、Taobao、RetailRocket)及一个真实市场(Thumbtack)上验证:连续时间衰减机制已接近充分,其他条件化信号冗余或有害(公共数据上统计不显著,最高0.06 NLL;真实市场中至多轻微有害)。我们并非首次发现衰减有效,而是提供将‘条件化无效’转化为可验证、可认证结论的工具,并坦诚披露分析中遇到的读出/泄露陷阱及其修正。

原文摘要 · Abstract (English)

Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities. But does any of it improve the timing, and how would you know? A null ("feature X doesn't help") is only meaningful if the model could have found a signal. We make two contributions--a method and a measurement--to answer this credibly. (i) A screen-and-confirm protocol that certifies whether a candidate signal improves a TPP's event-timing likelihood: a positive control plants a coupling of known strength and confirms the model recovers it, so a real-data null can be read as "no signal" rather than "weak method." The control is validated for categorical and continuous encodings, and on a real clock-driven dataset (NYC taxi hour-of-day). (ii) A model-free ceiling quantifying how little of customer-return timing is point-predictable at all (a single-digit percentage of gap variance from any covariate; returns are near-memoryless). With these we certify a clean result on three public benchmarks (Amazon, Taobao, RetailRocket) and a real marketplace (Thumbtack): the inter-event clock--continuous-time decay, long known to beat frozen-intensity models--is nearly sufficient, and the conditioning the field keeps adding is redundant or harmful on top of it (statistically null on the public benchmarks, at most 0.06 NLL; null to mildly harmful on the marketplace). We do not claim to discover that decay helps; our contribution is the tools that turn "conditioning doesn't help" into a checkable, certified statement--plus an honest-evaluation account of the read-out/leakage pitfalls we hit and retracted.

客户行为时间建模模型验证衰减机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。