Imagine this illustrative scenario: three of your forty branches lose half as many customers as the rest. Everyone has a theory. The regional lead says it is the managers. The managers say it is the customers. Head office wants "the winning routine" rolled out by next quarter.
Before you write that playbook, find out what those three branches actually do, and whether it would survive the trip to branch number four. That is top performer analysis: define an outcome, pick the people who have genuinely achieved it, reconstruct how they work, and choose one practice to test somewhere else. Done well, it produces something specific enough to try and honest enough to argue with.
The method has a research pedigree. Positive-deviance studies in healthcare examine exceptional performers to generate hypotheses, then test those hypotheses more widely. Bradley and colleagues laid out that sequence, and their design is the backbone of this guide. It is a reference for how to investigate, not evidence that any particular commercial practice works.1
The worksheets below are working tools, not a validated scoring system. Copy them, fill them in, and change what does not fit.
Decide what "top" means before you pick anyone
Start with records before leaders nominate their favourites. Reputation, likeability, or one good quarter can stand in for sustained results when the eligibility rule is vague.
Start from a decision rather than a curiosity. "We need to choose which repair-handover practice to test across our branches" gives the study a narrower job than "Let's find out what great managers do." Then write down what counts as strong performance, and go and look at the records behind it. Reliable performance measurement is a prerequisite in the positive-deviance approach; without it you are studying reputation.1
Check the boring things: the same definitions everywhere, the same reporting period, the same denominator, no branch quietly booking returns as new jobs. If the records cannot carry the selection, say so, and describe your participants as experienced operators rather than verified top performers. A playbook built on a flattering measure travels badly.
| Field | Complete before recruiting |
|---|---|
| Decision | Which operating practice might we adopt or test? |
| Outcome | What result matters, and how is it calculated? |
| Evidence | Which records support it, and who checked them? |
| Period | What span includes ordinary work and the busy periods that matter? |
| Eligibility | Which roles, locations, and work types are comparable? |
| Context | What differences in resources or customers must be recorded? |
| Guardrail | What must not get worse while chasing the main outcome? |
| Uncertainty | What could make the apparent performance misleading? |
Pair the headline measure with a guardrail. Repair speed with repeat repairs. Sales with refunds. Ticket closure with reopened tickets. A branch that "closes tickets fast" is not excellent until you know what "closed" means there.
Pick cases for what they can teach you, including the ones that failed
Choose participants because their experience can answer the question, not because they are available. Purposeful sampling selects cases for the information they can provide about the thing under investigation.2
Build a recruitment table with the verified measure, the context, the reason for inclusion, and the participation status. Then add contrasting cases: a comparable branch with ordinary results, and one that tried a promising practice and abandoned it. The abandoned case can reveal costs or operating conditions that a successful adopter's account leaves unexplained.
Confront selection bias in writing. Who does the definition exclude? Who declined? Whose work is missing from the records? Do not quietly swap an unavailable top performer for whoever picks up the phone. Review recruitment against the important situations still unexplained, not a round number. If your budget ends first, record the gaps and narrow the claims.
Ask for last Tuesday, not for their philosophy
Use the same core topics in every interview and leave room for surprises. GOV.UK's guidance on in-depth interviews recommends open, neutral questions, real examples, and follow-up prompts.3
Do not open with "Your handover process is clearly why you perform so well." You have just handed them the answer. Ask for the last time instead:
- "Take me through the last job that needed a handover."
- "What did you receive, and what did you do first?"
- "What told you it was ready for the next person?"
- "Where did you have to stop, check, or ask for help?"
- "What document, screen, or message were you using?"
- "Describe a recent time when that approach did not work."
- "What would another branch need before trying this?"
Ask to see things where you can: a blank checklist, a redacted job record, a dummy case walked through on screen. Keep the interview on the workflow and away from customer and employee details you do not need.
Then label what you heard honestly. An account of work is an account, not an observation. Memory is imperfect, which is why contextual inquiry combines questioning with watching people work in their own environment.4 Mark each piece of evidence as reported, demonstrated, documented, or observed. A manager who describes a daily review but cannot show one has still given you an idea. It just carries a label. For more on getting past the first answer, see our customer interview follow-up guide.
Turn stories into rows
After each interview, write one row per distinct practice. A practice is an action with a trigger and an owner. "Good communication" is a mood. "The receiving technician confirms the next action before accepting the handover" is something a second branch could try on Monday.
| Practice | Evidence | Context | Challenge | Decision |
|---|---|---|---|---|
| Trigger, actor, and action | Interview reference; record or observation reference | Required resources, workload, and authority | Counterexample, missing evidence, or competing explanation | Investigate, test, adapt, or set aside |
Keep performance evidence in its own column or sheet. A job record can verify that a checklist was used. It cannot, on its own, tell you the checklist produced the results.
When you group rows, compare actions rather than phrases. Two managers who both say "we take ownership" may be describing completely different routines. Write down the differences before deciding they are one practice.
A worked example
Everything in this example is fictional: the business, the branches, the records, the observations.
An equipment-repair network wants fewer jobs bounced back because the original fault was never fixed. The operations lead checks how branches classify repeat repairs first, and finds that one branch records returns as new jobs. That branch drops out of the comparison until its records are reconciled. Selection first, interviews second.
She then interviews eligible branch managers and technicians about recent handovers. The goal is to choose one test, not to announce a network-wide standard.
| Candidate practice | Fictional evidence | Context and challenge | Proposed treatment |
|---|---|---|---|
| Receiving technician confirms the unresolved symptom before accepting a job | Riverside's manager describes it; a redacted handover form shows the field | Harbour has a similar field, and staff admit skipping it when the queue grows | Observe whether confirmation actually happens; investigate staffing and ownership |
| An unresolved-diagnosis tray with a named reviewer | Meadow's technician demonstrates it with a dummy job | Found at one branch; depends on a senior reviewer other branches lack | Keep as a promising rare practice; test whether a different reviewer can run it |
| A morning team meeting | Mentioned at every branch interviewed | Content varies; ordinary-performing branches hold them too | Do not treat meeting frequency as the explanation |
Notice what happens to the first candidate. The form is not the practice. The useful hypothesis is that an explicit confirmation by the receiving person surfaces missing information before the job moves on. Harbour has the form and skips it, which is exactly the kind of counterexample that sharpens a finding.
The tray survives too. One branch doing something is thin evidence, not disqualifying evidence. The open question is whether the responsibility transfers without Meadow's particular expert.
What the team rejects is "copy Riverside's whole routine." It proposes one bounded handover test and keeps the tray for further investigation.
Count, then argue with yourself
Record how many interviewed cases support each practice, and next to it how many had a real opportunity to use it. "Not mentioned" and "asked and not used" are different answers. Neither a vivid story nor a high count should settle the decision on its own.
Then try to knock your own finding down. Could experienced staff, simpler jobs, or better parts availability explain the result? Does the practice also exist at branches with ordinary outcomes? HM Treasury's evaluation guidance is blunt about examining alternative explanations before crediting a cause.5
Write the shortlist rationale in four parts: evidence quality, plausible mechanism, operating conditions, and transfer difficulty. Keep uncommon practices visible when their evidence deserves it. And resist the urge to turn interview counts into rankings of cause or estimates of how common a practice is across the industry. They are counts from the people you talked to.
Test one practice before you standardise it
Hand the receiving team a practice card, not a policy:
- Trigger: when a job moves between technicians.
- Action: the receiving technician confirms the unresolved symptom and the next action.
- Prerequisites: access to the job record and authority to ask for clarification.
- Evidence to collect: completed handovers, missing-information incidents, repeat repairs, added handling time.
- Stop condition: workload or service-quality problems that mean pausing the trial.
- Review decision: keep, revise, test further, or drop.
Agree the review period and the decision rule before the trial starts. Check whether the practice was actually followed, and note anything else that changed at the same time. Where you can, keep comparable work running under the existing process alongside it. A simple before-and-after improvement does not rule out the other explanations you listed.5
The first test may only tell you whether the practice is usable under another branch's conditions. That is a real result. Label it as one.
Finish each playbook entry with its evidence references, the situations it applies to, the exceptions, an owner, and the trigger for the next review. Our evidence-backed insight guide is a useful check that the recommendation stays inside what you actually learned.
Where PaidInsight fits
PaidInsight is still in pre-launch development.6 Its Top Performer Analysis workflow is built for this kind of study: you define the qualification, select operators, and send invitations. The workflow supports individual phone interviews with Penny, our AI interviewer, and follow-up questions aimed at concrete examples and process details, within the interview's time and question limits.7
When you are ready, choose Generate report and confirm the interview set to request analysis. The analysis compares reported practices across the qualified interviews included from that set, with supporting interview evidence and counts. The report workflow can turn that material into an evidence-linked playbook with a designed PDF export.8
Those inclusion checks do not independently verify an operator's performance; that remains your records and your selection. Individual calls also give you a different format from a group discussion. And a count across your interviews is still a count, not proof of cause. The test at the end is yours to run.