Benchmarks
Proving a Lift
How to tell whether dunning changes improved recovery, using comparable baselines, enough volume, and segmentation to distinguish lift from noise.
This page is about dunning recovery, the campaign that starts when a subscription payment fails. Lift on a cancel flow is a save rate question with the same logic and no cohort wait; see Save Rate.
Prerequisites: Comparing Periods, Tools, and Migrations and Natural Variance. This page assumes both.
Complete cohorts. Every outcome counted.
What incremental lift is
Incremental lift is the change in recovery rate caused by a change to your recovery process, relative to a baseline measured before the change. It requires a rolling analysis baseline captured before any optimization, or pre-change daily cohorts, ideally with all four outcomes broken out.
Lift is the change in recovery rate relative to a baseline measured the same way. It requires a baseline. It is not claimed without one.
The change can be a new tool or a new setting, such as campaign length or email cadence. The method is the same for each.
Why one number proves nothing
A recovery rate after a change, on its own, isn't evidence of anything. A typical account swings 20-30 points between its best and worst rolling window over a multi-year history, with no change to the process, as Natural Variance shows. Without a baseline range, an improvement is indistinguishable from variance, and any claimed lift is unfalsifiable.
Two things make most claimed lifts unprovable: not enough volume, and a customer mix shift that segmentation would have exposed.
What a lift measurement requires
A recovery lift is only provable against a baseline captured before the change goes live. One of two things has to be true.
1. You captured a baseline before the change. Run rolling analysis on your existing recovery, on complete daily cohorts with all four outcomes, over enough history to see your variance band, before you apply any optimization. Native recovery reporting often stops when you switch off the native flow, so this has to happen deliberately and in advance. Comparing Periods, Tools, and Migrations covers the frozen baseline you get when it doesn't.
2. You have full pre-change data. Daily cohorts with all four outcomes broken out is ideal. The floor is successful payments and total churn, with cancellations included in churn. Anything thinner can't support a defensible lift claim.
Either way, a baseline is only comparable when it uses the same methodology as the after.
When the baseline isn't ideal
Most teams don't have daily cohorts. They have a dashboard number from a formula they've never seen. What you have falls into one of four tiers.
- Best case. Daily cohorts with all four outcomes. Build the baseline from them directly.
- Good enough. A reported rate for a clean historical period, plus confirmation of what its formula includes and excludes. You know the bias, so you can adjust for it or at least name it.
- Minimum. A historical rate of unknown formula. Treat any comparison against it as directional only, not as proof of a lift, and write down the gap.
- Insufficient. A number with no known method, no clean date range, or overlap with the transition. Don't accept it as a baseline.
Volume decides what you can detect
A lift is only measurable when the account has enough completed campaigns for its rolling band to be narrower than the change you're trying to detect. Low volume accounts have wide bands. An actual improvement can sit inside the noise for a long time, and a claimed improvement can be pure variance.
On a small account, use the largest rolling window your completed days will support, and expect to wait longer than 60 days before a difference is distinguishable from noise.
The failure mode is a lift that was never measurable because the volume never supported it. Check your band before the change, so you know in advance whether the improvement you're hoping for is large enough to show up.
Segment before you attribute
The biggest false lifts and false declines come from customer mix shifts, not process changes. A promotion or viral moment two to three months ago puts a wave of first and second renewal customers into recovery now. They fail more and recover less, so the blended rate drops with no change to the process. The reverse happens when the surge ages out, and if that follows your change, it looks like a lift. Mix shifts like these explain a surprising share of natural variance.
Segmentation is how you narrow that variance into something actionable. Same data, sharper question.
Before attributing any movement to the change, split first renewal customers from the rest, or compare long tenured customers alone across the before and after. If the segmented rates are flat and the blended rate moved, the story is mix, not process.
The procedure
A lift uses the same before and after method as Comparing Periods, Tools, and Migrations, condensed to four steps.
- Establish the baseline and its formula. Capture it before the change, meeting the requirements above. See what the baseline counts.
- Identify a clean measurement period. Exclude the first N days after the change, where N is the previous campaign length. Then use complete cohorts only, in rolling windows, with 60+ days of completed cohorts after the transition. See the transition rule.
- Reconcile the formulas. If the two sides differ, name the difference and quantify it. A baseline that excluded cancellations comes in roughly 10-20 points high, which can make an improvement look like a decline. See reconciling formulas.
- Account for natural variance. A 2-3 point difference inside the account's natural variance band is not a meaningful performance change in either direction. See testing against your band.
Three conclusions are possible: the change improved recovery, performance is flat within variance, or it declined.
What an estimate is and isn't
A lightly optimized setup can leave measurable recovery work unfinished. Treat any expected upside as an estimate until you measure it against a baseline with the method above.
How lift gets faked
Three measurement errors produce a lift that isn't there.
- A baseline fixed at the worst period. Pin the baseline to one low month and every later upward swing counts as improvement, including gains in the business itself.
- Seasonal highs credited to the change. A change that goes live ahead of a naturally strong stretch collects the credit for the season.
- No variance range. Without the high and low of the account's history, nothing separates the change from noise, so every swing above the baseline passes as lift.
All three are forms of item 11 on Practices This Methodology Rejects. Claiming a lift with no baseline, or one measured a different way, is item 5.
Don't measure it this way
- Claiming lift without a pre-change baseline. There's nothing to check the claim against. Capture the baseline first.
- Comparing to a baseline computed with a different formula. The difference measures the methodology gap, not the change. Use the same formula on both sides.
- Crediting a blended rate move to the change without segmenting. A mix shift moves the blended rate too. Split out first renewal customers first.
- Calling a within band difference a lift. Check the band first.
- Using a baseline that overlaps the transition. Campaigns from both configurations are in it. Move the range back.
- Expecting a small account to prove a lift on 60 days. Its band is too wide. Use the largest window the data supports and wait longer.
The full list is on Practices This Methodology Rejects.
Prerequisites: Comparing Periods, Tools, and Migrations, which owns the before and after method this page applies; Natural Variance, which sets the band a lift has to clear.
Previous: Realistic Recovery Rate Range.
This is the last page in the library. Before any comparison, the two pages worth rereading are the Glossary of Recovery Measurement Terms and Recovery Rate: the Formula.