Build the Benchmark from Your Best Customers Playbook 2026
Every other play compares accounts with a benchmark. This is how to build it: pull the customers who renewed flat or up, shrank or left, strike the unfair matches, compare usage and revenue by customer type and stage, label thin segments low confidence, and have your CS lead approve what good looks like – then check it every quarter.
Your best customers already show you what good looks like: how many of their people use the product each week, which features they rely on, and what they pay at each stage. Written down, that's a benchmark every account can be measured against. The health score backtest is the case for building one; this page is how.
The play sorts every customer with enough history by what happened at renewal (renewed flat or up, shrank, left), compares usage and revenue by customer type and stage, and has your team approve the result. It's the first step of every other play in this library, starting with the weekly check.
Measure it on segments with a checked benchmark, accounts the weekly check covers, and whether the benchmark still separates renewed from left.
How it works11 steps
01SignalPull every customer who has had time to grow or leave
One row per customer who has come up for renewal after at least 12 months paid, or who left in the last 24. Four things per row:
- Tenure & stage: start date, and the three windows you'll measure: Trial, New (first 90 days paid) and Established (the 12 weeks ending 90 days before the scheduled renewal, the same window for every account, so leavers aren't measured while they wind down).
- Usage by feature, per window: weekly active users, seats, and weeks each key feature was used.
- ARR & plan history: ARR now (or at leaving), ARR 12 months earlier, plan changes.
- Outcome, at the renewal: Grew (renewed with ARR up 10% or more, or upgraded plans with ARR at least unchanged), Renewed flat (renewed at the same or higher ARR, short of Grew), Shrank (renewed at lower ARR), Left (canceled). Defaults to tune. Customers who renewed at the same or higher revenue (Grew and Renewed flat) are the benchmark cohort; Shrank and Left stay separate.
A user is active in a week when they take one of your key actions – the ones that deliver the value, not a login. Under about a hundred accounts, do all of this by hand in the spreadsheet below.
02Human stepYour CS lead picks the fair matches
Before comparing anything, the CS lead strikes customers who would skew it and writes the reason in Fair match note: internal, partner or test accounts; deep discounts or custom contracts; accounts migrated from another product; tracking gaps of two weeks or more; seasonal customers measured out of season. Keep the struck list – someone will ask.
03ScoreCompare usage and revenue by customer type and stage
Split by what changes what good looks like: plan tier or pricing model first, then size (seats or ARR band). Start with no more than four segments. Within each segment and stage, take the median of each measure for each outcome group, with its count beside it – medians, because one large account drags an average.
The median of customers who renewed at the same or higher revenue is the benchmark. The grew median sits beside it, with its own count, as the growth line for expansion plays; the shrank and left medians show what a gap looks like. Use the benchmark as a reference, not a risk threshold: about half of its own cohort sits below it.
Use matched peers for descriptive comparisons. Before a measure triggers a risk flag, check its threshold against later renewal outcomes; for expansion plays, check it against later expansions.
Usage and revenue go together. Say established mid-market customers who renewed at the same or higher revenue have 62% of seats active weekly, use four key features and pay $24,000; those who left had 21%, one feature and $15,000. The $9,000 between them is a comparison, not a forecast. (The numbers are illustrative.)
04DecisionCheck there are enough customers to compare
Rule, default: a segment and stage needs at least 10 fair-match customers who renewed at the same or higher revenue. Fewer, and it isn't benchmarked yet: a median from a handful of customers isn't a benchmark. Show the counts beside every benchmark. Few or no cancellations doesn't block a descriptive comparison, but it limits how far you can test a risk threshold.
05ActionLeave thin segments out, and write down why
For each account in a segment below that count, set Recommended move to No move, with the reason in Move reason and a Review date at the next quarterly refresh. The weekly check skips them rather than guessing. If two small segments look alike, merge them and recount.
06DecisionDecide how far to trust it
Default: 20 or more who renewed at the same or higher revenue, with 12 months of usage tracking behind them, is Provisional. Fewer, or less tracking history (say you changed analytics tools in the spring), is Low confidence. A benchmark becomes Checked only after the frozen version is tested on later renewals it wasn't built from, with its misses, false alarms and uncertainty written down.
07ActionDraft the benchmark: what you saw, apart from what you think
One row per segment, stage and measure that passed. We keep what we saw apart from what we think might be true and still need to test: What we saw (customers who renewed at the same or higher revenue used reports four weeks in four) goes in the benchmark; What we think (reports cause the growth) goes on its own tab with how you'll test it.
08ActionMark thin segments low confidence, and say so
When there is little history, we trust the benchmark less, and we say so. Set Benchmark confidence to Low; every move the weekly check drafts in that segment carries the low-confidence note below.
09Human stepYour CS lead approves what good looks like
The CS lead, one CSM or AM, and someone from sales or finance for the revenue side, on the agenda below. Expect an argument about at least one measure – that's the benchmark earning trust. Nothing reaches an account record until the CS lead signs it off.
10ActionWrite the benchmark onto each account record
Set Benchmark segment, Customer stage and Benchmark confidence on each company in HubSpot. From here, the weekly check compares each account with its own segment and stage.
11OutcomeCheck it every quarter
Before you change anything, test the frozen benchmark on the renewals that came due last quarter: did the gaps show up on the accounts that went on to leave or shrink, and not on the ones that renewed at the same or higher revenue? Count the misses (leavers with no gap) and the false alarms (renewers with one), and write down how sure you can be with that many renewals. Record the answer in the Refresh log: Held up (keep it; after its first pass, mark it Checked), Measure dropped (one missed – drop it), or Rebuilt (it didn't separate – rebuild the segment). Only then re-label outcomes and re-run the medians. A "what we think" row that keeps showing up is still an association, not proof that the feature caused growth; it stays on the To test tab until you've tested it. Fold in what the eight-week review found about which moves worked.
Setup and templates
Sources: usage (a saved query in PostHog, Amplitude or Mixpanel, run per window), billing (Stripe or your invoices) for ARR and plan changes, the CRM for owner and notes. Account-level numbers need your analytics tool's account add-on (group analytics in PostHog, Accounts in Amplitude, Group Analytics in Mixpanel) and a group call in your app. Save the query, then subscribe to it (in Amplitude, add it to a dashboard and subscribe to that) so the result lands in your inbox or Slack on the schedule you set. By hand: export all three to one spreadsheet.
The spreadsheet: tab Customers (Company · Segment · Stage window · Weekly active % · Key features used · Seats · ARR now · ARR 12 months ago · Outcome · Fair match · Fair match note) · tab Benchmark (Segment · Stage · Measure · Benchmark median · Grew median · Shrank median · Left median · Counts · Confidence) · tab To test (Segment · What we think · How we'll test it) · tab Refresh log (Quarter · Segment · Result · Misses · False alarms · What changed · Signed off by).
Company properties in HubSpot (Starter or above; free HubSpot allows 10 custom properties in total), or your CRM's account object:
| Property | Type | Filled by |
|---|---|---|
| Benchmark segment | Dropdown select: your segments | CS lead, at approval |
| Customer stage | Dropdown select: Trial · New · Established | Weekly check |
| Customer outcome | Dropdown select: Grew · Renewed flat · Shrank · Left · Too soon | Quarterly refresh |
| Fair match | Dropdown select: Yes · No · Not reviewed | CS lead |
| Fair match note | Single-line text | CS lead |
| Benchmark confidence | Dropdown select: Checked · Provisional · Low · Not benchmarked | Quarterly refresh |
| Recommended move | Dropdown select: each play's moves · No move | Weekly check (this play: No move only) |
| Move reason | Single-line text | Weekly check |
| Review date | Date picker | Quarterly refresh, then weekly check |
Success measure and Review result belong to the plays that run moves; see the eight-week review.
The automation here uses HubSpot workflows (any Professional hub). On Free or Starter, run the same saved view and create the tasks by hand.
The saved view is Benchmark to review: Fair match is Not reviewed, or Benchmark confidence is Low or Not benchmarked. Each quarter the CS lead gets one task, "Benchmark refresh: [quarter]", with the spreadsheet linked. The workflow waits until a week before the approval meeting, then creates it due that day.
Fair match note: "Struck – [discount / custom contract / migrated / tracking gap / test account / out of season]. [initials], [date]."
Low-confidence note (on each move):
Low confidence – [segment] has [n] comparison customers and [months] of tracking. Check this gap by hand before acting. Recheck [date].
Peer lines in customer messages: only send a "teams like yours usually" line your benchmark backs; otherwise say what you saw on their account.
Approval meeting agenda (45 minutes):
- Segments and counts, after the struck customers (5)
- The struck list – anyone we shouldn't have removed? (10)
- Renewed flat or up vs shrank vs left, per stage, and the growth line: which measures made the cut (15)
- Confidence labels and segments left out (5)
- "What we think" rows and how we'll test each (5)
- Sign-off, and the next refresh date (5)
No move: "No move – [segment] not benchmarked: [n] comparison customers, fewer than [minimum]. Recheck [date]."
How Accoil fits
Accoil is a customer-success consultancy with our own tooling, and this is a play we run with you, not for you. We build the benchmark with your team from your own successful customers, and you approve it. Every week we check every account in the agreed segment against it, and the recommended move and its evidence land on the account record in your CRM – we recommend HubSpot – at a pace your team can actually run. Your team makes every call. Nobody from Accoil contacts your customers. After the agreed window we review what worked together. We also review the benchmark, more often in the first 90 days, less after that, and you keep the playbook, the benchmark and the analysis.
Running it yourself? Swap HubSpot for Salesforce or Pipedrive. The play stays the same. On Salesforce, the field writes need API access (Enterprise and up, or an add-on on Professional). On Pipedrive, automations and date triggers need the Growth plan.
Thirty minutes with Kate, Simon or Peter on one example segment: what a successful customer looks like, where the gaps are, and whether there's work worth doing this quarter.
Talk to a founder →Keep reading
Seat & Feature-Adoption Upsell Playbook 2026
Accounts show what they'd pay more for by how they use what they have. This play compares each account in a segment every week with the customers who bought your add-on or added seats, checks the records before calling it a gap, and puts one move on the account in your CRM: an add-on deal, a seat true-up – or no move, with the reason noted.
The Eight-Week Review Playbook 2026
Eight weeks after a move, this play brings the account back up and asks what changed. It keeps activity, progress and business result apart, checks what else played a part, and puts the verdict on the CRM record: it worked, change it, stop it, or give it more time – with the reason noted. Each quarter, what worked goes back into the benchmark.
Monday-Morning Account Triage Playbook 2026
Your best customers show you what good looks like. Every Monday, this play compares each account in a segment with that benchmark, ranks what each gap is worth, caps the list at what your team can run, and puts the week's moves on the account records in your CRM: close a gap, grow the upside, or leave it alone with the reason noted.
Every playbook, one download
All 32 workflows as print-ready playbooks — diagrams included. Plus every new workflow as we publish it.