RevOpsSuccess

Health Score Backtest Playbook 2026

Put the health score you already have beside what happened at last year's renewals: what it caught, what left while green, which red flags renewed fine, how much warning it gave. If it held up, keep it and retest next year. If it fell short, rebuild on a benchmark of what your successful customers did, picked on earlier renewals and tested on later ones.

Peter Preston · Co-founder, Accoil·Updated Sep 2026·Advanced
Measure it onLost ARR that was green 90 days outRed flags that renewed fineDays of warning before renewalGrowth accounts the score flagged as risk

Your renewals from last year are already an answer key. Every account either grew, stayed or left, and your health score said something about each of them 90 days before. Put the two side by side and you find out what the score is worth and, if it falls short, where to look instead: what your successful customers did.

This play runs that test once a year: keep the score and write down why, or rebuild on a benchmark of what your successful customers did and pay.

Measure it on the share of lost ARR that was green 90 days out, the share of red flags that renewed fine, the days of warning, and – the business result – whether next year's losses show up in time to act.

How it works11 steps

01SignalSet last year's scores beside what happened
Accoil

One row per account whose renewal came up in the last 12 months:

  • Score 90 days out: the score as it stood 90 days before the renewal date (default), not today's. Today's has already absorbed the result.
  • Customer outcome: Grew (ARR up 10% or more, or upgraded plans with ARR at least unchanged), Left (canceled, or ARR down 25% or more), Stayed (everyone else who renewed). These are backtest buckets; the benchmark cohort includes only completed renewals at the same or higher revenue.
  • ARR before and after from billing or the CRM.
  • Usage before renewal: weekly active users and key actions over the six months ending 90 days before renewal, from PostHog, Amplitude or Mixpanel. Nothing after that date goes into the score or the benchmark.

If your score is a number, use the green, amber and red cut-offs your team already works to. If you never kept score history, there's nothing to test yet: start a weekly snapshot now and go straight to the benchmark. Under about a hundred accounts, do it all by hand in a spreadsheet.

EmitsScore 90 days outGrew, stayed or leftARR before and afterUsage up to 90 days out
02ScoreCount what the score caught, missed and got wrong
Accoil

Call an account flagged if it was amber or red 90 days out. Exclude accounts that had already given notice 90 days out; the score gets no credit for catching those. Mark each with one Backtest result:

  • Caught: flagged, then left.
  • Missed while green: green, then left.
  • Grew while flagged: flagged, then grew – the score read growth as risk.
  • Flagged, then stayed: record separately whether an intervention (a call, a save plan, an escalation) is documented before the renewal. Report those two groups apart; neither tells you whether the flag was right or the move saved the account.
  • Right, no flag: green, then stayed or grew.

Then four numbers: the share of lost ARR that was green, the share of flagged accounts that left, the median days of warning, and the share of growth accounts flagged. For days of warning, take every lost account and count the days from the start of its last unbroken run of amber or red to its renewal, from the weekly score history. A lost account still green at renewal counts as zero; report accounts with missing history apart. Say 140 renewals: 46% of lost ARR was green, one flagged account in three left, the median warning was 34 days, and nine of 22 growth accounts were amber. (The numbers are illustrative.)

EmitsLeft while greenFlagged, then stayedDays of warningGrew while flagged
03ActionPut each miss on its account record
HubSpot

Write every Missed while green, Flagged, then stayed and Grew while flagged onto the company in HubSpot, with the owner's task below.

04Human stepThe account owner notes what the score didn't see

Two lines in What the score missed: what they saw at the time, and what in the product or the contract would have shown it. Before counting a miss against the score, check the records:

  • Leaves usage can't show – acquired, closed, lost the budget. Exclude them, and count them apart.
  • Gaps in score history – if more than one account in five has no score 90 days out (default), write "low confidence" beside the verdict.
EmitsOwner's noteLeaves usage can't showGaps in score history
05DecisionDecide whether the score earned its keep

Rule, defaults to tune, set before you look: the score fell short if more than 20% of lost ARR was green 90 days out, or fewer than one flagged account in three left, or the median warning is under 60 days. Otherwise it held up.

06ActionIf it held up, write down why and retest next year

Record No move in the backtest review with the four numbers and a retest date a year out. A written pass gives the team a reason to trust red.

07ScoreRebuild the benchmark from what successful customers did
Accoil

Don't reweight the inputs – replace the question. A score asks "does this account look healthy?" A benchmark asks "is this account doing what our customers who stayed and grew did, at this stage, at this price?"

Take the customers who renewed at the same or higher revenue, compare what they did in the product and what they pay with those who left, and keep only the measures that separate them. Pick the measures on earlier renewals. Test them on a later period only if it holds enough renewals and losses to judge both rules. Otherwise build on the history you have, freeze the rule, and record both scores at 90 days out from now on; review quarterly, and keep collecting outcomes while the result is too close to call. Add the owners' notes on the misses as candidate measures. The benchmark play has the steps.

EmitsRenewed flat or upTheir product usageWhat they pay
08ActionTest the benchmark on later renewals

Run the test on the later renewals, and score the old health score on the same renewals: 90 days before each renewal, was the account short of the benchmark? Replay both rules week by week, using only what was known that week. Short counts as flagged; on or ahead counts as green. Count the same four numbers. Recommend switching if the benchmark leaves less lost ARR unflagged and gives more warning, without doubling the flagged accounts that stayed (default). It also shows what a score can't: accounts that stayed but sat well short of their peers – the growth a risk flag never looks for.

09Human stepThe CS lead and RevOps decide whether to switch

With the two results side by side, they switch, keep the score, or run both for a quarter – recorded in the backtest review.

10ActionRun the weekly check on the benchmark instead
HubSpot

The weekly check compares every account in the segment with the benchmark and puts the recommended move on the HubSpot record. Retire the score from views and tasks so the team works from one list.

11OutcomeReview it a quarter on, then retest every year

On the review date, answer three questions: are owners acting on the moves (activity), are gaps closing (progress), and did any account leave without being flagged first (the result)? Note what else played a part, then change it, stop it, or give it time – the eight-week review has the method. Run this backtest again next year, with the benchmark in the score's seat.

Setup and templates

Sources: score history from your CS platform (Gainsight's scorecard snapshots or ChurnZero's ChurnScore history) or the weekly exports you've saved; renewals and ARR from billing (Stripe) or the CRM; usage from a saved query in PostHog, Amplitude or Mixpanel; reasons for leaving from CRM notes. Account-level numbers need your analytics tool's account add-on (group analytics in PostHog, Accounts in Amplitude, Group Analytics in Mixpanel) and a group call in your app. Save the query, then subscribe to it (in Amplitude, add it to a dashboard and subscribe to that) so the result lands in your inbox or Slack on the schedule you set. By hand: one spreadsheet row per renewal, with the properties below as columns.

Company properties in HubSpot (Starter or above; free HubSpot allows 10 custom properties in total), or your CRM's account object:

PropertyTypeFilled by
Score 90 days outDropdown select: Green · Amber · Red · No historyBacktest
Customer outcomeDropdown select: Grew · Stayed · LeftBilling or CRM
Backtest resultDropdown select: Caught · Missed while green · Flagged, then stayed · Grew while flagged · Right, no flag · ExcludedBacktest
What the score missedMulti-line textAccount owner
Intervention documentedSingle checkboxBacktest
Benchmark saidDropdown select: Short of benchmark · On benchmark · AheadReplay

The automation here uses HubSpot workflows (any Professional hub). On Free or Starter, run the same saved view and create the tasks by hand. The saved view is Backtest misses: Backtest result is Missed while green, Flagged, then stayed or Grew while flagged, and What the score missed is unknown. Each creates a task for the owner, "Backtest: what did the score miss at [company]?", due in two weeks:

[company] was [green / red / amber] on [date], 90 days before renewal, and [left / stayed / grew] ([ARR before] to [ARR after]). What did you see at the time, and what in the product or contract would have shown it? If no usage could have shown it, set Backtest result to Excluded and say why.

The backtest review (a shared doc, one per year, linked from the saved view):

FieldValues
Recommended moveRebuild the benchmark · No move
Move reasonText
Success measureText
Review dateDate
Review resultSwitched · Running both · Kept the score · Stopped

Backtest, [year]: [n] renewals, [n] excluded. Lost ARR green 90 days out: [x]%. Flagged accounts that left: [x]%. Median warning: [n] days. Growth accounts flagged as risk: [x]%. Benchmark on the test period: [same four]; old score on the same period: [same four]. What owners said the score missed: [top three]. Decision: [move], because [reason]. Success measure: [measure]. Review [date].

No move: "No move – the score held up: [x]% of lost ARR green, [n] days of warning. Retest [date]."

How Accoil fits

Accoil is a customer-success consultancy with our own tooling, and this is a play we run with you, not for you. If the backtest is included in our agreed work, we compare your historical scores with actual renewal outcomes using the records you provide. We build the benchmark from your own successful customers with your team, and you approve it. Your team decides whether to keep the score, test the benchmark alongside it, or switch. Every week we check every account in the agreed segment against it, and the recommended move and its evidence land on the account record in your CRM – we recommend HubSpot – at a pace your team can actually run. Your team makes every call. Nobody from Accoil contacts your customers. After the agreed window we review what changed together, and you keep the playbook, the benchmark and the analysis.

Running it yourself? Swap HubSpot for Salesforce or Pipedrive. The play stays the same. On Salesforce, the field writes need API access (Enterprise and up, or an add-on on Professional). On Pipedrive, automations and date triggers need the Growth plan.

Talk to a founder

Thirty minutes with Kate, Simon or Peter on one example segment: what a successful customer looks like, where the gaps are, and whether there's work worth doing this quarter.

Talk to a founder →
Share this playbookLinkedInX

Keep reading

YOUR EXISTING EVENT STREAMAPION BENCHMARKWATCHGAPYESLEAVE ALONESIGNAL · ACCOILWeekly check: renewals in 90days vs customers who renewedRenewal date & stageUsage vs renewersARR vs renewed peersChampion still active?SCORE · ACCOILCheck the records before callingit a gapGap & what it's worthRecords complete?Tickets & billingBudget or org changesDECISIONWhere does it sit against therenewers?DECISIONRoom to grow at renewal?ACTION · HUBSPOTDraft an expansion look for therenewal callACTION · HUBSPOTNo move: note why on the recordfor the next checkACTION · HUBSPOTDraft a value review on theHubSpot recordACTION · HUBSPOTDraft a save plan for the AM andCSM, with evidenceHUMAN STEPAM or CSM decides and runs themoveOUTCOMEEight weeks on: is the gapclosing? Then: did it renew?
Account ManagementSuccess

Renewal-Risk Radar Playbook 2026

Renewals go best when the account already looks like your customers who renewed. From 90 days out, this play compares every renewing account with them each week – usage and revenue – checks the records, and puts one move on the account in your CRM: an expansion look, a value review, a save plan with the CSM – or no move, with the reason noted.

Starter
YOUR EXISTING EVENT STREAMAPIYESLEAVE ALONEHIGH STAKESTHE RESTSIGNAL · ACCOILWeekly check: the gap to thebenchmark widensGap vs a month agoPeer features droppedARR & plan vs peersRenewal 90+ days outSCORE · ACCOILCheck the records before callingit a gapSeasonal or data blip?Champion still active?Open tickets & talksBudget or org changesDECISIONIs there a move worth makingnow?DECISIONWhich move fits this account?ACTION · HUBSPOTDraft a save conversation withthe evidenceACTION · HUBSPOTAutomated lane: re-engagementemails, logged on the recordACTION · HUBSPOTNo move: note why on the recordfor the next checkHUMAN STEPCSM or AM decides and runs theconversationOUTCOMEEight weeks on: did use comeback? What else played a part?
SuccessAccount Management

Churn-Risk Save Playbook 2026

Your customers who renewed show you what steady use looks like. This play checks every account against that benchmark each week, catches the ones drifting further away while the gap is still worth closing, and puts the move on the account in your CRM: a save conversation, a re-engagement sequence your team chose to automate – or no move, with the reason noted.

Starter
YOUR EXISTING EVENT STREAMAPIYESLEAVE ALONEOVER 90DINSIDE 90DRIGHT-SIZESIGNAL · ACCOILWeekly check: fewer seats in usethan peers your sizeActive vs paid seatsPeers' active-seat %Plan & ARR vs peersActivity per seatSCORE · ACCOILCheck the records before callingit a gapSustained, notseasonal?Quiet vs never startedDays to renewalCuts, champion &ticketsDECISIONIs there a move worth makingnow?DECISIONWhich move fits this account?ACTION · HUBSPOTDraft a seat review with bothpaths on the recordACTION · HUBSPOTDraft a reactivation plan for theunused seatsACTION · HUBSPOTDraft a right-size quote with theseat listACTION · HUBSPOTNo move: note why on the recordfor the next checkHUMAN STEPAM or CSM decides and runs themoveACTION · APPCUES Guide never-started seats to afirst result in-appACTION · STRIPEIssue the right-size quote at thestandard rateOUTCOMEEight weeks on: seats in use, or aright-size agreed?
Account ManagementRevOps

The Shelfware Confession Playbook 2026

Seats paid and seats in use drift apart. This play compares each account's seat use every week with customers its size that renewed at the same or higher revenue, checks the records before calling it a gap, and puts one move on the account in your CRM: a seat review with both paths, a reactivation plan, a right-size quote – or no move, with the reason noted.

Intermediate
The playbook pack

Every playbook, one download

All 32 workflows as print-ready playbooks — diagrams included. Plus every new workflow as we publish it.

On-page playbooks stay ungated. Downloading subscribes you to new workflows from Accoil – unsubscribe anytime.