Analyze
Comparing Recovery Rates Across Periods, Tools, and Migrations
How to compare a dunning recovery rate before and after a settings change or a tool switch without fooling yourself: consistent formulas, complete cohorts, the transition rule, and the natural variance test.
This page is about dunning recovery, the campaign that starts when a subscription payment fails. Cancel flows follow different rules.
Switching dunning tools, changing recovery settings, or reviewing performance over time all come down to comparing a before and an after. A team switches tools, compares two numbers, sees one it doesn't like, and starts making changes, when the data never said what it seemed to say.
1. Capture the baseline before you switch
Native recovery analytics usually exist only while the native flow is running. Take the pre-switch number deliberately, through an export or a read-only connection, or it's gone, and you end up comparing an in-progress cohort on the new tool against a frozen number on the old one. That comparison makes a healthy migration look like a regression.
The best baseline is daily cohort data with all four outcomes broken out, since you can rebuild any rate from it. If the old tool only gives you a headline number, take the number and write down what it counts.
2. Confirm what the baseline counts
A baseline is only comparable when it uses the same methodology as the after. Complete cohorts only, with no in-progress campaigns in the denominator. All four outcomes counted. A clean date range with no overlap with the transition. Enough volume to smooth natural variance.
Different tools define and calculate recovery rate differently. Some exclude cancellations from the denominator. Some count a recovery in the period it occurred rather than the period the payment failed. Even when two tools use the same formula on paper, the underlying data can differ: which charges the tool includes, how it assigns campaign start dates, whether it counts stopped or manually resolved campaigns.
If you're comparing performance across tools, you need a consistent data source and a consistent calculation for both time periods. Dashboard numbers from two different platforms are not comparable.
Most people can't get this from a dashboard. If that's what you have, confirm what the number includes and excludes, treat the comparison as directional rather than precise, and never accept a baseline that overlaps the transition.
Four traps recur in native dashboards.
The baseline disappears when you switch the native flow off.
A monthly cohort dashboard checked mid-period comes in low by construction. Its campaigns haven't resolved. Establish when the cohort closes instead of arguing the number.
Check what the denominator counts before comparing two retention numbers. A retention figure that keeps long inactive subscribers in the active bucket flatters itself.
Some dashboard numbers are snapshots, not measurements over the selected date range. A count of subscriptions currently in dunning is the state of the system, and it doesn't change when you move the date picker. Compare it across two date ranges, or against a windowed number from another tool, and you're comparing a stock to a flow. Translate the snapshot into a windowed figure first, campaigns started or resolved in the period, then compare.
3. Exclude the transition
When you migrate between tools or change recovery settings such as campaign length, retry timing, or email cadence, there's a transition period where data from both configurations is active at once. Including this data contaminates both your before and your after.
Identify the transition period and exclude it entirely. For a migration, exclude the first N days after go-live, where N equals the previous tool's campaign length. Campaigns already in progress under the old tool carry over and finish under different logic. If the old campaigns ran 14 days, exclude the first 14 days on the new tool. If 30, exclude 30. Then wait for 60+ days of completed cohorts before comparing.
4. Measure the after on complete cohorts, in rolling windows
While a campaign is running, its outcome is unknown. Drop the active campaigns from the denominator and you've removed the campaigns whose outcome you don't know yet. Keep them in the denominator and count them as unrecovered and you've called them lost before they resolved. In the worked example on Recovery Rate: the Formula, the same day produces 80% one way and 40% the other, and neither is a rate.
Exclude entire time periods that still have active campaigns. Only analyze complete cohorts where every campaign has reached a final outcome. How a cohort completes is on Daily Cohorts and Complete Cohorts.
Sum the complete cohort days into rolling windows rather than calendar months; Rolling Analysis covers how. The longest rolling window your completed days support is the cleanest single number for the after.
Comparing "last 3 months" to "last 6 months" isn't a comparison when the 6 contains the 3. The two numbers share most of their data, so any difference comes out muted and looks like flat or declining performance. Use non-overlapping periods, the prior 3 months against the most recent 3, and exclude the transition so it sits in neither.
5. Reconcile the two formulas
If both sides count all four outcomes on complete cohorts, compare directly. If they don't, name the difference and quantify it where you can. The most common difference is a baseline that excludes cancellations, which comes in roughly 10-20 points high depending on the cancellation rate.
6. Test the difference against your band
Recovery rate fluctuates when nothing about your process has changed. If your 30-day rolling rate over the past year falls between 62% and 74%, a month at 64% isn't a crisis. Natural Variance is the size of that fluctuation.
Establish your baseline range before drawing conclusions about any individual period. A 2-3 point difference inside your band is not a signal in either direction. Only a move outside the range, sustained across rolling windows, is evidence that something changed.
7. Rule out the other explanations
Did the mix of outcomes shift? Every campaign ends in one of four outcomes: Card Update, Successful Retry, Cancellation, or Passive Churn. Your rate can hold at 68% while card updates fall from 35% to 25% and retries rise from 33% to 43%. Chart each outcome over time.
Did anything else change? Everything that affects whether your customers want to stay subscribed moves your recovery rate. Seasonality. Holidays, end of year, and summer shift recovery behavior, and from late November through December more cards fail and it's harder to get subscribers' attention. Acquisition surges. A major promotion or a viral moment two to three months ago puts a wave of first and second renewal subscribers into recovery now. Product or pricing changes. A price increase, a new subscription offering, or a redesigned website shifts retention behavior, which surfaces in passive churn. Broader economic conditions. Macro factors move failure rates and recovery behavior across your whole customer base. Before attributing a change in recovery rate to your dunning process, ask whether anything else changed during the period.
Did the customer mix shift? Subscribers failing on their second charge recover at much lower rates than long tenured customers, so if your mix shifts toward new customers, your overall rate drops even when each segment's performance holds. Segment your data, even separating first renewal customers from the rest, and you'll explain a surprising amount of the variance.
Was it a processing anomaly? A processor returns a higher volume of vague "declined" errors from a temporary system issue rather than from the cards themselves. Or a processor reprocesses a batch of charges from weeks or months ago, flooding your funnel with old failures that have low recovery probability. Check the distribution of charge errors for the period in question. An unusual concentration of one error type, or charges already attempted several times before entering your campaign, is a reason to investigate before changing strategy.
For the specific case of proving that a change produced a lift, use the same discipline: predefine the comparison, keep the periods clean, and separate signal from variance.
Two things that look wrong after a switch and aren't
The failed payment count goes up. A tool that retries more often logs more attempts, and every attempt that fails is a recorded failure. The share of payments that fail hasn't moved. Watch the recovery rate, not the failure count.
Churn drops for a few weeks, then jumps about a month in. Customers move into longer campaigns, so cancellations show up later than they used to. The early drop is churn deferred, not churn reduced, and the number normalizes once those campaigns end.
How long to wait
30+ days of completed cohorts before drawing a conclusion about your performance. 60+ days of completed cohorts, after the transition, before comparing a before and an after.
Don't measure it this way
- Comparing dashboard numbers from two tools. Different formulas, different data. Rebuild both sides from the same export.
- Comparing a resolved cohort against a mid-period rate. One has finished and one hasn't. Wait for the cohort.
- Accepting a baseline that overlaps the transition. Campaigns from both configurations are in it. Move the range back.
- Comparing nested windows. The longer one contains the shorter one. Use non-overlapping periods.
- Comparing a snapshot to a windowed metric. One is a stock, the other a flow. Translate the snapshot into campaigns started or resolved in the period first.
- Calling a within-band difference a result. Check the band first.
These rejected practices all have the same shape: they mix methods, periods, or unresolved cohorts and then ask the result to carry more meaning than it can.
Prerequisites: Rolling Analysis, which builds the windows both sides of a comparison are made of; Natural Variance, which sets the band a difference has to clear.