Two managers, one call, two different scores
Your ops manager listens to a call and gives the rep a 7. Your sales lead listens to the same call and gives it a 3.
Nobody is lying. They're scoring different things. The ops manager heard a friendly rep who booked the appointment. The sales lead heard a rep who never asked what the customer had already tried, quoted a price range over the phone, and let the customer pick the time slot.
Here's why it matters more than it sounds: the rep can't act on a 7 and a 3. When two managers disagree, the rep learns that scores are opinions and coaching is a mood. They stop changing behavior and start managing whoever reviewed them that week.
Calibration fixes that. Not by making everyone agree on everything — by making everyone agree on what the words on your scorecard mean.
First, figure out which kind of disagreement you have
Not all score gaps are the same problem. Before you run a meeting, sort the disagreement into one of three buckets.
1. Definition gaps. Both managers heard the same moment. They disagree on whether it counts. One thinks "confirmed budget" means the customer said a number. The other thinks it means the rep said a number and the customer didn't object. This is the most common gap and the easiest to fix.
2. Attention gaps. One manager missed something. They were skimming, or they stopped listening at minute four, or they scored from memory of the rep rather than from the call. Fixable with process, not debate.
3. Weighting gaps. Both managers heard the same thing and agree it happened. They disagree on how much it should cost. "He interrupted her twice" — is that a one-point deduction or a failed call? This one you have to decide as a business. There's no objective answer.
Most teams try to solve all three with a longer rubric. Longer rubrics make definition gaps worse, because there are more words to interpret.
Rewrite your scorecard items as things you can hear
The single biggest source of disagreement is scorecard language that describes a quality instead of a behavior.
"Built rapport." "Showed empathy." "Controlled the call." "Professional tone." Two reasonable people will never score those the same way, because they're describing an impression.
Replace each one with something a stopwatch or a transcript could settle.
| Vague item | Hearable version |
|---|---|
| Built rapport | Used the customer's name at least once after the greeting |
| Uncovered the need | Asked what the customer had already tried to fix it |
| Controlled the call | Asked the next question within 5 seconds of the customer finishing |
| Confirmed decision process | Asked who else would be involved in the decision |
| Set a firm next step | Got a verbal yes to a specific day, time, and who will be present |
Test each item with this question: could two people who dislike each other still agree on the score? If not, rewrite it.
Some items genuinely require judgment — "did the rep handle the price question well?" is a real skill and you can't reduce it to a keyword. For those, don't try to make the item objective. Make the anchors objective instead.
Anchor the scores with real clips, not adjectives
A 1–5 scale with descriptions like "exceeds expectations" is a coin flip. A 1–5 scale anchored to actual recordings is not.
Pick one call from your own library for each score level on your hardest rubric items. Save the timestamps. Now "a 4 on discovery" doesn't mean good, it means like the Reyes call at 2:15, or better.
"A 5 on price handling sounds like the Tuesday afternoon roof call — she asked what number he had in his head before she said anything about ours."
Managers can argue about adjectives forever. It's much harder to argue that this call is better than that clip.
Run a 30-minute calibration session
Do this monthly. Weekly for the first month if scores are all over the place.
Before the meeting: pick two calls. One clean, one messy — the messy one is where you learn. Every manager scores both independently and submits before the session. No comparing notes beforehand, or you get groupthink instead of calibration.
In the meeting:
- Put the scores side by side. Don't discuss the calls yet.
- Find the item with the widest spread. Start there. Skip the items where everyone agreed — you're not there to feel good.
- Play the specific 30 seconds in dispute. Not the whole call. The moment.
- Each manager says what they heard and what rule they applied. In that order. "I heard her say 'I'll email you some options.' I scored next step as a 1 because there's no date and no confirmation from the customer."
- Decide the rule out loud. Write it down as a one-line note on the scorecard item. "Email follow-up with no date = 1, even if the customer sounds interested."
- Move to the next widest gap. Stop at 30 minutes even if you're not done.
That written note is the actual output of the meeting. Not agreement — documentation. Six weeks from now, when a new manager scores the same situation, they read the note instead of guessing.
The rule that saves the most time
Make one person the tiebreaker and say so out loud. Usually the sales leader or the owner.
Calibration isn't a democracy. When two managers can't converge after playing the clip twice, the tiebreaker decides, it gets written down, and everyone scores that way going forward — including the manager who disagreed. Endless consensus-seeking is why calibration meetings get canceled.
Track your drift with a number
You don't need statistics for this. You need one number you check monthly.
Take the two calls everyone scored. For each rubric item, note the gap between the highest and lowest score. Average those gaps.
- Average gap under 0.5 points: you're calibrated. Spot-check quarterly.
- 0.5 to 1.0: normal. Keep the monthly session.
- Over 1.0: your scorecard is the problem, not your managers. Cut items, rewrite the vague ones.
One more check worth doing: have each manager rescore a call they already scored three weeks ago, without looking at their old score. If a manager disagrees with themselves by more than a point, the item is too fuzzy to be coachable. Fix the item.
What to tell the reps
Reps notice when scoring is inconsistent, and they talk about it. Get ahead of it.
Tell them plainly: "We found we were scoring discovery differently. We fixed the definition. Here's the new one, here's the clip that shows what a 4 sounds like, and last month's scores on that item don't count."
Throwing out the old scores costs you nothing and buys real credibility. Defending scores you know were inconsistent costs you the rep's attention for the next year.
Then hold the line. Once the definition is written down, coach to it the same way every week. Consistency is the whole point — a slightly imperfect rubric applied identically beats a perfect rubric applied three different ways.
The short version
- Sort disagreements into definition, attention, or weighting. Only weighting is a real judgment call.
- Rewrite every scorecard item so two people who don't like each other could still agree.
- Anchor score levels to actual recorded clips from your own team.
- Run a 30-minute monthly session on the two calls with the widest spread, and write the rule down.
- Name a tiebreaker.
- Track the average high-low gap. Over one point means fix the scorecard.
If you're scoring calls with software, the same discipline applies — the value of automated scoring is that it applies one definition every time. But somebody still has to decide what "confirmed the next step" means for your business. That decision is manager coaching work, and it's worth half an hour a month.