The playbook
- 1
Days 1–3 — Scoping
Pick one tight use case, define baseline metrics, agree pass/fail criteria upfront.
- 2
Days 4–9 — Build
Build the agent against the scoped workflow. Connect to real systems. Test on real data.
- 3
Days 10–13 — Live shadow run
AI runs in parallel with humans. All outputs reviewed. Variance measured against baseline.
- 4
Day 14 — Decision gate
Review metrics. Hit pass criteria → expand. Miss → adjust scope or stop. No emotional decisions.
What you walk away with
Hard data on AI vs human performance
Defendable ROI projection for full rollout
Clear go/no-go decision
Capped downside (£3–8k investment)
Frequently asked
Why not just go full?+
Pilots de-risk both sides. They prove the case to your stakeholders and stop us building the wrong thing.