Behavioural Change Toolkit · Plan

Pilot Design and Experiment Card

Purpose: find out you’re wrong while it’s still cheap.

A pilot that succeeds and changes nothing was a demonstration, not a pilot.


Avoiding the pilot trap

The classic failure: pilot with a volunteer team, an enthusiastic manager, unlimited support and daily attention. It succeeds. You scale it to teams with none of those things and it collapses — then everyone concludes the people were the difference.

Design against it:

Do Don’t
Pick an ordinary team Pick the most enthusiastic team
Pick a merely competent manager Pick your best manager
Give the support you can actually scale Give unlimited hand-holding
Run through a busy period Run through a quiet week
Include a sceptic Volunteers only

If a pilot needs three people on site every day to work, you have learned that the design doesn’t work — that’s a genuinely valuable finding, and much cheaper now than in month four.


Experiment card

Card

Experiment: _______________ Behaviour targeted (B#): _______________ Barrier addressed: _______________

Hypothesis: “If we _______________, then [who] will [do what], because [which barrier is removed].”

Where: _______________ How many people: _______________ Duration: _______________ (minimum 2 weeks — week 1 is novelty)

Primary measure: _______________ Baseline (captured before starting): _______________ Success threshold: _______________ Counter-measure (what we check hasn’t got worse): _______________

We will stop / change if: _______________ Decision date: _______________ Who decides: _______________

⚠️ Capture the baseline before you start. The single most common pilot error. Without it you cannot tell whether anything happened, and you will be reduced to arguing about impressions.


Comparison — pick the strongest you can afford

Design Strength Effort
Before/after, one team Weak — anything could explain it Low
Before/after + a comparison team Decent — controls for the general environment Low
Two teams, two different interventions Good — tells you which is better Medium
Staggered rollout across several teams Strong and natural in a wave plan Medium

The staggered rollout is usually the best value. You’re rolling out in waves anyway; sequencing them deliberately gives you a comparison group for free, at no extra cost and no ethical awkwardness.


What to collect

Quantitative (daily, automated where possible)

  • The behaviour count itself
  • Time taken
  • Error / rework rate
  • Counter-measure — the thing you’re worried about

Qualitative (twice weekly, 10 minutes)

  • What was annoying today?
  • What did you do instead when it didn’t work?
  • What would you change?

The qualitative data is where the design fixes come from. The quantitative data is what convinces the steering group. You need both, for different audiences.


Review — 45 minutes

  1. What did we predict? Read the hypothesis aloud, unchanged.
  2. What happened? Numbers first, no interpretation.
  3. What surprised us? ⭐ The most valuable question in the meeting.
  4. What did people do that we didn’t design for? Workarounds are design feedback, not misbehaviour.
  5. What are we changing? Must be at least one thing.
  6. Scale, adjust, or stop?

Outcome

Result: ☐ Scale ☐ Adjust and re-run ☐ Stop What we’re changing: _______________ Who told us (credit them publicly): _______________ What we still don’t know: _______________


Reporting a pilot that didn’t work

This is a test of your credibility, and an opportunity.

“The pilot didn’t get the adoption we expected, and it told us why: the process assumed people were at a desk, and 60% of them aren’t. That’s a design assumption we’d have carried into all six regions. Fixing it now costs two weeks. Finding it in October would have cost the programme.”

Framed this way, a failed pilot is the best possible advertisement for piloting — and for you. Leaders who have been burned by big-bang rollouts recognise the value immediately.


Measurement Plan — Leading and Lagging · Diffusion of Innovations · T — Experiment Note

← Back to the Behavioural Change Toolkit

Last updated: 11 August 2026