Calibration meetings exist because managers rate the same performance differently. The fix, as most performance guides describe it, is a room where those ratings get compared and pulled into line. What almost none of those guides mention is a January 2024 Harvard Business Review analysis arguing that the fix has its own failure mode: a calibration meeting rates each employee twice — once by their manager, once by the committee — and the second layer can compound bias as easily as remove it. That is not an argument against calibration. It is an argument against running it on good intentions alone.

Quick answer: Performance review calibration is a structured meeting where managers compare draft ratings for their teams against a shared rubric and shared evidence, before ratings are finalised, so that the same standard of performance earns the same rating regardless of which manager wrote it. Done well, it takes 60–90 minutes per group of 15–25 employees, follows a fixed presentation order, and ends with a documented reason for every rating that changed. Done badly, it becomes a forced distribution in disguise, a seniority contest, or a number that quietly reverts the moment the room empties.

A group of managers in a calibration meeting discussing draft performance ratings around a conference table

What Calibration Actually Changes (And What It Doesn't)

Calibration is not a second performance review. Nobody re-marks the employee's work from scratch. It is a comparison exercise: managers bring the ratings they have already drafted, plus the evidence behind them, and a facilitator walks the group through each one so the ratings can be tested against each other rather than sitting in isolation.

A calibration session has to answer four questions:

  • What does each rating actually mean in behavioural terms, not just a number?
  • Whose evidence supports this rating, and would it survive being read aloud?
  • Does this rating look consistent with others at the same level, doing comparable work?
  • If it changes, who owns writing down why — and where does that reason live afterwards?

Calibration is usually the specific agenda item inside a wider talent review meeting: the talent review decides what happens next for a group of people, while calibration is the piece that makes their ratings comparable in the first place. The distinction that gets lost most often is the one between calibration and forced ranking — frequently discussed as the same exercise with different names. They are not, and the difference matters both operationally and legally.

Comparison diagram showing calibration as discussion-based against forced ranking as quota-based

Calibration starts from evidence and asks whether ratings are consistent. Forced ranking (also called stack ranking) starts from a quota — say, 15% of a team must be rated below average — and adjusts ratings to fit it, regardless of whether performance actually varied that much. The practical risk is a facilitator under time pressure treating the two as interchangeable: "we need someone in the bottom band" is a forced-ranking instinct wearing a calibration meeting's clothes.

The Bias Calibration Is Supposed to Remove — And the Bias It Can Add

The reason calibration exists is that unmanaged ratings drift in predictable, nameable ways. A useful calibration facilitator can name these on sight, because naming a bias out loud in the room is often enough to interrupt it.

Four rater biases calibration is meant to catch: leniency and severity, recency, halo/horn effect, and similar-to-me bias

The scale of the problem behind these four biases is larger than most performance guides let on. A widely cited study by Scullen, Mount and Goff, published in the Journal of Applied Psychology in 2000, decomposed performance ratings from more than 4,000 managers rated by multiple bosses, peers and direct reports. Idiosyncratic rater effects — essentially, which rater happened to be holding the pen — accounted for 53–62% of the variance in ratings across the two samples studied. The ratee's actual measured performance accounted for only 21–25%. Put plainly: more of a typical rating's variation traces back to who wrote it than to what the person actually did.

That is the strongest case for calibration that exists. It is also, per the HBR analysis above, the reason calibration can misfire: subjecting a rating to a second round of group judgement adds a new source of variance rather than removing the first one, unless the process itself is disciplined. Three failure patterns recur:

  • Anchoring on the first name discussed. Whoever gets presented first in a rating band sets the reference point every later case gets compared to, regardless of whether that first case was actually representative.
  • Seniority or airtime dominance. The most senior person in the room, or simply the person who talks the most confidently, pulls the group's judgement toward their view — independent of whose evidence is stronger.
  • Groupthink under time pressure. With twenty people to get through in ninety minutes, disagreement gets smoothed over rather than resolved, and the room defaults to whatever the manager originally proposed.

None of this is an argument for skipping calibration — the untreated problem is worse. It is an argument for treating the session as a process that needs designing, not a meeting that runs itself once the ratings are on a slide.

Running the Session

The five-step version below is deliberately mechanical. Calibration sessions fail less often from a bad rubric than from a good rubric applied inconsistently under time pressure, so the fix is procedural as much as it is philosophical.

Five steps for running a calibration session: set the rubric, randomise the order, challenge with evidence, document the rationale, lock and communicate

1. Set the rubric in behavioural terms, before anyone sees a name

Agree, in writing, what separates each rating band — not "exceeds expectations" as a label, but the specific evidence that earns it. Circulate this before the session, not during it, so the argument in the room is about evidence against a fixed bar, not about what the bar should have been.

2. Randomise the presentation order

Never work through a list alphabetically, by seniority, or in the order managers happen to submit it. A randomised order is a direct, low-cost countermeasure to anchoring: it stops whichever case comes up first from quietly setting the standard for everyone who follows.

3. Let the least senior voice go first on any challenge

When a rating is questioned, ask the newest or most junior manager in the room for their read before the most senior one speaks. This single change is one of the more effective, least-discussed countermeasures to seniority dominance, because once the most senior person has stated a view, disagreement becomes socially expensive.

4. Move a rating only when someone can cite the evidence

"I just don't see them as a 4" is not a reason to change a rating. "Here is the deliverable that was late three times this quarter, and here is what the rubric says about reliability at that level" is. If nobody in the room can produce evidence, the original rating stands.

5. Document the rationale, not just the new number

Every rating that moves in the room needs a written sentence attached to the employee's record explaining why, who raised it, and what evidence was cited. This is the single most commonly skipped step, and it is the one that determines whether the session was fair or merely felt fair to the people in the room.

Where the Good Intentions Fall Apart

Most calibration guidance stops at "hold the meeting." The meeting is rarely where it goes wrong.

The rating quietly reverts afterwards

A rating gets calibrated up or down in the room, and three weeks later the manager delivers a different number in the 1:1 — sometimes because they disagreed with the room and never said so, sometimes because the calibrated number never made it back into the system of record. Almost no published guidance addresses this directly, which is striking given that it is the easiest way for a calibration session to have been theatre. The fix is structural: the calibrated rating should write back to the employee's record automatically, with the rationale attached, so there is no gap for a quiet revert.

The employee was never in the room

Calibration decides an outcome for someone with no visibility into the discussion and, usually, no idea it happened at all. If a rating changes, the employee is owed the reason — not the transcript, but the specific evidence cited. Treating this as an afterthought is how a rigorous process still ends up feeling arbitrary from the receiving end.

Small teams have nothing to calibrate against

A manager with one or two direct reports cannot meaningfully compare them against each other. Widen the comparison group to a function or level across the business instead of a single reporting line, and lean more on peer and self-input as a substitute.

Cross-department calibration compares incomparable jobs

A "4" for a finance analyst and a "4" for a field engineer are not obviously the same standard. The workable pattern is two passes: calibrate within a function first, then run a senior-only pass across functions focused only on genuine outliers — not a re-litigation of every rating.

Potential gets calibrated as if it were performance

Performance measures what someone did. Potential is a judgement about what they might do next, and it is far more exposed to similar-to-me bias because it rests on projection, not evidence. Calibrate the two as separate exercises, ideally on separate fields, so a strong performance rating cannot silently inflate a potential rating it has no bearing on — the same distinction behind a well-run 9-box exercise.

The Legal Question Nobody Answers Precisely

"Is a forced distribution legal?" comes up constantly, and most answers are vague. Here is what is actually on the record — none of it legal advice, and a specific process should be checked with employment counsel.

JurisdictionWhat has actually been testedThe practical implication
United StatesFord Motor Company paid $10.5 million in 2001 to settle age- and gender-discrimination claims after its forced-curve "ABC" rating system disproportionately assigned its lowest grade to older and male employees; Ford scrapped the system within six months of the suit being filed. Microsoft faced comparable claims over its stack-ranking system before retiring it in 2013.A forced distribution is not automatically unlawful, but a curve that produces a statistically skewed outcome against a protected group is exactly the pattern that has previously supported a disparate-impact claim.
United States (2025 development)Following an April 2025 executive order, the EEOC began administratively closing disparate-impact-only charges, including those brought under the Age Discrimination in Employment Act. The underlying legal theory has not been repealed by Congress or the courts.Reduced federal enforcement is not the same as reduced legal exposure. A forced curve that skews by age or another protected characteristic remains a litigation risk through private claims, even where the EEOC itself is less likely to pursue it.
United KingdomNorman v Lidl Great Britain Ltd (2024) was a redundancy case, not a calibration case, but it is the closest and most instructive UK precedent: the tribunal found that using a degree qualification as a redundancy-selection criterion was indirect age discrimination, because employees over 60 were statistically less likely to hold one, and Lidl had no justification on record for the criterion.The same logic applies to any selection matrix, including one fed by calibrated performance ratings: a criterion or a rating pattern with a disparate impact on a protected group needs an objective justification on file before it is used, not after a claim is lodged.

The thread connecting both jurisdictions is documentation. A calibration process that keeps a written rationale for every changed rating is not just fairer to run — it is the record that would need to exist if a rating pattern were ever challenged.

A Session Design That Holds Up

Put together, the fixes above compress into a structure that scales from one department to a full review cycle.

StageWho's involvedLengthOutput
Rubric circulationHR/People team1 week beforeBehavioural rating definitions agreed and shared
Draft ratings submittedLine managers3 days beforeRatings + supporting evidence per employee
Within-function calibrationPeer managers + facilitator60–90 min per 15–25 peopleRatings tested, outliers challenged, changes documented
Cross-function passSenior leaders + facilitator30–45 minGenuine cross-team outliers reviewed only
Write-back and communicationHR/People teamWithin 48 hoursFinal ratings and rationale attached to each record; employees told what changed and why, where applicable

The 48-hour write-back window matters more than it looks: the longer that gap runs, the more room there is for a rating to quietly drift back to what the manager originally proposed.

How StaffCircle Supports Fair, Documented Calibration

Every failure mode above has the same root cause: the calibration decision and the system of record are two different things, connected only by someone remembering to update a spreadsheet.

One rating, one record, no reversion window

Performance ratings, goal progress and competency assessments already sit in a single record per employee. A rating changed in calibration updates that record directly, rather than a spreadsheet a manager has to remember to reconcile afterwards — closing the gap where a rating can quietly revert.

A manager writing a documented rationale for a calibration decision at a desk

Every changed rating keeps its rationale

A rationale, an author and a timestamp can be attached to any rating change made during calibration, so the record shows not just the final number but why it moved and who raised it — the documentation gap that most calibration processes leave open, and the one a challenged rating pattern would need to answer.

Performance and potential stay on separate fields

Because performance ratings, 9-box placement and succession data live as distinct fields on the same platform rather than one blended score, a strong performance rating cannot silently inflate a potential rating during calibration — the two get discussed, and changed, independently.

These capabilities are described as currently built; verify the exact behaviour in a live demo before relying on them for a specific compliance requirement.

Final Thoughts

Calibration is not a fairness guarantee. A badly run session can add bias as easily as an unmanaged one, and the legal record shows what happens when a distribution-driven process cannot explain itself after the fact. What calibration reliably does is give an organisation the chance to catch the roughly half of rating variance that traces back to who happened to write it, provided the session is designed with the same rigour as the rubric it is meant to enforce.

That means a fixed rubric circulated in advance, a randomised order, a rule that ratings move only on cited evidence, and a documented rationale for every change — written back to the record within 48 hours, not left to survive the walk back to someone's desk. Book a demo to see how StaffCircle keeps calibration decisions, rationale and the employee record in one place.

FAQ

What is performance review calibration?

A meeting where managers compare draft ratings and the evidence behind them, so the same standard of performance earns the same rating regardless of who wrote it.

What is the difference between calibration and forced ranking?

Calibration adjusts ratings based on evidence, with no fixed quota. Forced ranking sets a quota first and adjusts ratings to fit it. In practice, a facilitator under time pressure can blur the two by treating "someone needs to be in the bottom band" as a valid reason to move a rating.

Is a forced distribution or bell curve for performance ratings legal?

Not automatically unlawful, but it carries real litigation history: Ford paid $10.5 million in 2001 after its forced curve disproportionately rated older and male employees at the bottom band. This is not legal advice; check a specific policy with employment counsel.

How long should a calibration meeting take?

60–90 minutes for 15–25 employees within a function, plus a shorter 30–45 minute cross-functional pass for genuine outliers. Stretching one session to cover 50+ people is where most calibration meetings lose rigour.

Who should attend a performance calibration session?

Peer managers rating comparable roles, plus an HR or People facilitator. Senior leaders join only the cross-functional pass, to review flagged outliers, not to re-run every rating.

What happens if managers disagree during calibration?

A rating moves only if someone can cite specific evidence against the agreed rubric. Asking the most junior manager in the room to speak first, before the most senior one, reduces the tendency to simply defer to seniority.

Does calibration actually reduce bias, or can it add more?

Both, depending on how it's run. A 2024 Harvard Business Review analysis argues that rating someone twice — manager, then committee — can introduce new bias through anchoring, seniority dominance and groupthink, unless the session is structured to guard against them.

What is rating inflation, and does calibration fix it?

The tendency for ratings to drift upward, often to avoid a difficult conversation. Calibration catches it when ratings are compared against an evidence-based rubric, but not if the rubric is vague enough to justify almost any rating.

What's the difference between calibrating performance and calibrating potential?

Performance calibration compares evidence of what someone did. Potential calibration is a judgement about a future, different role, and is more exposed to similar-to-me bias because it rests on projection rather than a finished body of work. Discuss and document the two separately.

How do you calibrate ratings for a small team with only one or two people per manager?

Comparing one or two people against a distribution built from unrelated teams is close to meaningless. Widen the comparison group to a function or level across the business, and lean more on peer and self-input.

Should distributed or remote teams calibrate differently?

Same mechanics, but the evidence has to work harder: a facilitator has less incidental visibility into a remote employee's work, so evidence attached to the rating before the session matters more.

What should happen after a calibration meeting ends?

Every changed rating should write back to the employee's record within 48 hours, with the rationale and evidence attached. The longer that gap runs, the more likely a rating drifts back to what the manager originally proposed.

Can an employee's rating be changed in calibration without them being in the room?

Yes — which is exactly why documenting the rationale matters. If a rating moves, the employee is owed the specific reason and evidence, not just a different number on their review.


About the author

Mark Seemann is the CEO and Founder of StaffCircle, the AI performance management platform for mid-sized organisations. He writes about performance management, employee development and the practical use of AI in HR. Connect with Mark on .