Skip to main content

Review cycles

How to run a performance calibration session

Calibration is where a set of individual manager judgements becomes one shared standard. Here is how to prepare the evidence, facilitate the meeting and close the loop afterwards.

By the CLEAR Talent team7 min read

Performance review cycle overview in CLEAR Talent

Key takeaways

  • Calibration moderates the standard, not the manager. The objective is comparability, not consensus.
  • Circulate the evidence before the meeting. A session that opens with reading has already lost half its time.
  • Work the outliers first: the highest ratings, the lowest, and anyone who has moved more than one band.
  • Every changed rating needs a recorded reason and a manager willing to explain it to the employee.

Calibration is the meeting where a pile of individual manager judgements is turned into one shared standard. Done well, it is the most effective control an organisation has over rating quality. Done badly, it becomes a negotiation in which the most confident manager wins.

The mechanics are not complicated, but they are easy to get wrong: the wrong people in the room, evidence that arrives too late to be read, or a distribution curve quietly treated as a quota. What follows is what to do before, during and after the session — and the failure modes worth naming out loud before you start.

What a calibration session is for

A calibration session brings managers together to compare how they have applied a rating scale across a group of employees. The subject of the discussion is the standard, not the person: does “exceeds expectations” mean the same thing in engineering as it does in customer support?

That distinction sets the boundary of the meeting’s authority. Calibration can legitimately challenge whether a rating is consistent with the evidence and with how comparable people have been rated. It should not be used to relitigate a manager’s day-to-day judgement about someone they work with daily and the rest of the room does not.

It is also not a ranking exercise. A distribution curve can be a useful diagnostic — if four in five people in a department land in the top band, something is wrong with either the standard or the scale — but a distribution used as a target turns calibration into an allocation problem and costs you the managers’ trust in the process.

Why manager ratings drift apart

Variation between managers is normal and mostly harmless. The variation worth correcting comes from a handful of recurring patterns.

Different internal standards
Two managers can read the same rating descriptor and picture very different levels of performance. Without a shared reference point — behavioural indicators, worked examples — the descriptor does most of its work in the reader’s imagination.
Recency
The last eight weeks are vivid and the first eight months are not. Ratings written without contemporaneous notes tend to describe the most recent quarter and call it a year.
Leniency and severity
Some managers rate high to protect their team; some rate low to look rigorous. Both distort comparability, and both become obvious the moment their distributions are placed side by side.
Team context
A solid performer in an exceptional team gets rated down by comparison, and a middling performer in a weak team gets rated up. The scale should measure the person against the standard, not against whoever happens to sit beside them.
Visibility
People whose work is presented to leadership are easier to rate highly than people whose work is only visible inside their own team. Ask what evidence exists for the quiet contributors before the ratings are fixed.

Before the meeting: prepare the evidence

Most calibration sessions are lost before anyone sits down. Five things need to be settled in advance.

  1. Fix the population and the scale

    Decide who is in scope and confirm that everyone is being rated on the same scale, with the same descriptors, for the same period. Mixed scales make the comparison meaningless.

  2. Publish the standard, not just the labels

    Circulate the rating descriptors with behavioural examples for each level. If your competency framework already defines indicators per level, use those rather than inventing new language for the meeting.

  3. Assemble the evidence per person

    Each rating should arrive with the material behind it: goal outcomes, competency ratings and the evidence attached to them, and any multi-rater feedback. A rating with no supporting evidence is a position, not an assessment.

  4. Send the distribution in advance

    Managers should see the group picture before the meeting — ratings by manager, by team, by level. People arrive better prepared to defend or revise a rating once they have seen where it sits.

  5. Agree the ground rules

    State who chairs, how long each case gets, how disagreements are resolved and who holds the final decision. Ambiguity on that last point is what turns calibration into a negotiation.

If the evidence pack cannot be read in twenty minutes, it is too long. If it can be read in two, it is too thin.

Running the session

Open by restating the purpose and the scale in a minute or less. Then start with the outliers rather than working alphabetically: the highest ratings, the lowest ratings, and anyone whose rating has moved more than one band since the previous cycle. That is where inconsistency lives, and those are the discussions that deserve the time.

  1. Ask for evidence, not adjectives

    “Reliable”, “a great team player” and “a real asset” are not evidence. The question to fall back on every time is simple: what did this person do, and what was the outcome?

  2. Test the rating from both directions

    Two questions do most of the work. “What would this person have had to do to be rated one band higher?” and “Would you give the same rating if they sat in a different team?”

  3. Time-box every case

    Give each person the same starting allowance and let the chair extend it deliberately. Without a clock, the first five names consume the session and the rest get rubber-stamped.

  4. Record the decision as you go

    Capture the outcome and the reason at the moment they are agreed. Reconstructing a rationale a week later produces a summary nobody in the room can stand behind.

Close with a different question. Not “are we happy with these ratings?” but “would this set of ratings look defensible to someone who was not in the room?” That is the standard the decisions will eventually be held to.

After the session: close the loop

A rating that changes in calibration has to be explained by the manager who owns the relationship, in their own words, with the reasoning that persuaded the room. Employees can accept a moderated rating. What they will not accept is a rating that changed for reasons nobody will articulate.

Keep the record — who was in the room, what changed and why. That trail is what makes the cycle defensible under internal audit, works-council review or a later dispute, and it is what lets you compare this cycle against the last one.

Then feed the outcome forward. A gap the meeting identified is development content: if three managers agreed that someone sits a band below the standard on a specific competency, that belongs in a development plan, not only in a rating field.

Finally, review the meeting itself. How much of the session went on evidence and how much on advocacy? Which managers consistently arrive with the strongest case? That is a training signal for the next cycle.

Five ways calibration goes wrong

The distribution becomes a quota
A guide curve used as a diagnostic is healthy. The same curve enforced as an allocation turns the meeting into a trading floor and teaches managers to inflate their opening ratings so they have something to concede.
Nobody brought evidence
Without goal outcomes, competency evidence and feedback in the room, the session rewards the most fluent speaker rather than the strongest case.
HR owns the outcome
HR should design and facilitate the process. Once HR is also making the rating decisions, managers stop owning them — and stop defending them to their teams.
It happens too late
Calibration scheduled after ratings have been communicated is theatre. It has to sit between the manager’s draft and the final rating, while the decision is still live.
The loop is never closed
If nothing visibly changes for employees or managers as a result, the next cycle gets treated as an administrative step to be completed rather than a decision to be made.

What tooling can and cannot fix

No platform will tell you what “exceeds expectations” should mean in your organisation. That is a leadership decision, and it is the one that determines whether calibration works at all. What software can do is remove the practical reasons a session runs badly.

Evidence in one place
In CLEAR Talent, goal scores roll up from Strategic Linked Goals automatically and competency ratings carry the behavioural evidence behind them, so the pack managers read is the same data the rating was made from.
Outliers surfaced before the meeting
Calibration views use heatmaps, outlier flags and forced-distribution simulations to show where ratings are inconsistent across managers — before anything is final, and with forced distribution as an optional guardrail rather than a requirement.
Multi-rater input, structured
360° feedback is collected against the same competencies as the review and anonymised by rater group, so peer input can be discussed in the room without exposing individuals.
A record that survives the cycle
Every change is logged with who, what and when, and one-click PDF and Excel exports produce both the pack for the meeting and the record afterwards.

Ava, the platform’s AI assistant, can draft the summaries and assemble the pack. It does not make the call. The rating decision stays with the managers in the room, which is the only place it can defensibly sit.

See how this works in practice

Book a walkthrough focused on your review cycle, your goal structure, and the decisions your managers actually have to make.