Quick answer: Proficiency levels are the graded scale — usually 4 to 7 tiers — that a competency framework uses to describe how well someone needs to perform a skill or behaviour in a given role. Managers score consistently against them only when each level is written as an observable behaviour rather than an adjective, the number of levels is chosen on purpose rather than copied from a template, and ratings go through a calibration session before they're finalised. Miss any one of those three and the scale still exists on paper, but two managers scoring the same evidence will land on different numbers.

The gap usually shows up at the first calibration meeting a company ever runs. Someone pulls ten "meets expectations" ratings from different managers and lays them side by side, and it becomes obvious that "meets expectations" meant five different things to five different people. Nobody was being dishonest. The scale never told them what a 3 looked like, so each manager filled the gap with their own judgement.

That is fixable without redesigning the whole framework. It just requires being deliberate about three things most frameworks leave vague: how many levels to use, what each one actually says, and what happens to a rating before it counts.

A competence record showing role standard, assessment, evidence attached and expiry tracked as the output of well-defined proficiency levels

What Are Proficiency Levels in a Competency Framework?

Proficiency levels are the rungs on the ladder a competency framework uses to describe capability. Where a competency defines what good looks like for a skill or behaviour — "commercial awareness", "line management", "SQL" — the proficiency scale defines how much of it a role needs, and gives an assessor language to describe where a real person currently sits. If you need the fundamentals of a competency framework itself first, our companion guide covers that ground.

A scale that actually works answers four questions for anyone using it:

  • What can someone at this level do, in observable terms, that someone at the level below cannot?
  • Would two different managers, looking at the same piece of work, place it at the same level?
  • Does the wording point at behaviour and output, or at personality traits nobody can evidence?
  • Is there a defined process for what happens when two assessments of the same person disagree?

Most frameworks answer the first question reasonably well and the other three not at all. That is where consistency breaks down.

Proficiency Level vs Skill Level vs Competency Level

In practice these terms are used interchangeably, and there is no industry-standard distinction between them. What matters is being consistent within your own framework. Some organisations reserve "skill level" for hard, teachable capabilities (a language, a tool, a qualification) and "competency" or "behavioural level" for judgement-based ones (communication, decision-making, leadership) — worth adopting even informally, because the two often need to be scored differently.

How Many Levels Should You Use? What SFIA, NHS KSF and the Civil Service Actually Show

Ask five HR consultants how many levels a framework needs and most will say "four or five", citing the Dreyfus model of skill acquisition. It's worth being precise about what that model says before borrowing it. Stuart and Hubert Dreyfus developed it in 1980 at UC Berkeley to describe how an individual's own competence develops over years — novice, advanced beginner, competence, proficiency, expertise, and sometimes a sixth stage, mastery — not as a scale for a manager to score someone against once a year. Citing it to justify a five-point annual scale borrows the stage names without the substance, which is why so many "5-level" frameworks read the same way.

Real, published frameworks used to assess people at scale don't converge on one number. Each picks a level count for a reason that's visible once you look at what the scale is for.

Table comparing SFIA, the NHS Knowledge and Skills Framework, Civil Service Success Profiles and a typical vendor scale by level count, seniority grading and published anchors
Level count is a design decision every published standard makes on purpose — not a rule of thumb.

SFIA, the Skills Framework for the Information Age, uses seven levels of responsibility — Follow, Assist, Apply, Enable, Ensure/Advise, Initiate/Influence, and Set Strategy/Inspire/Mobilise — because it grades autonomy and organisational influence across an entire career, and seven is what it takes to keep that range meaningful.

The NHS Knowledge and Skills Framework goes the other way: every dimension, including its core dimension of Communication, is scored on exactly four levels. Run across hundreds of thousands of staff, fewer, more clearly separated levels produce more consistent ratings than a finer scale would.

The UK Civil Service's Success Profiles Behaviours takes a third approach almost no vendor content mentions: the same nine behaviours — Seeing the Bigger Picture, Making Effective Decisions, Leadership and others — are described differently at six seniority bands, from AA/AO up to Director General. The behaviour doesn't change; the bar for meeting it does. That's the right model to reach for whenever one competency needs to mean something different for a junior starter and a department head.

A Decision Framework, Not a Default

Rather than picking a number and writing descriptors to fit it, work backwards from how the scale will be used:

  • Hard, teachable skills (a tool, a language, a certification) can usually support more levels, since progress is observable in discrete steps — closer to SFIA's seven.
  • Behavioural competencies (communication, judgement, leadership) are harder to separate reliably past four or five — closer to the NHS KSF's four.
  • Roles spanning a wide seniority range under one competency suit the Civil Service approach: keep the competency fixed, vary the expected bar by grade.
  • The bigger and more distributed the manager population, the fewer levels you can reliably hold consistent. Four levels rated well beats seven rated badly.

Make the choice on purpose, write it down, and be able to explain it later.

How to Write Proficiency Levels Managers Can Actually Use

Five-step process for writing proficiency levels: start from behaviour, choose the level count on purpose, write anchors not adjectives, validate with the people doing the job, pilot before rolling out

Step 1: Start From Behaviour, Not Personality

Write down what a person does and produces, not what kind of person they are. "Is proactive" cannot be evidenced by anyone; "raises risks to the project lead before they hit the deadline" can, because you can point to an instance of it happening or not.

Step 2: Choose the Level Count on Purpose

Use the decision framework above rather than defaulting to five because that's what the last template used. Write the reason down — "four levels because this applies across 40 sites and consistency matters more than granularity" — so it can be revisited deliberately rather than by accident.

Step 3: Write Anchors, Not Adjectives

This is the highest-leverage step, and the one most frameworks skip. A Behaviorally Anchored Rating Scale means every level carries a real example of what it looks like, not a synonym for "good". The NHS KSF's Communication dimension is a genuinely well-built example in the wild:

The NHS Knowledge and Skills Framework's four Communication levels, from day-to-day matters at level 1 to complex situations at level 4, each with real behavioural anchors
Each NHS KSF level carries its own bulleted behavioural indicators — not an adjective in sight.

Notice what it doesn't do: it never says level 4 is "excellent" or level 1 is "basic". It says level 1 communicates with a limited range of people on day-to-day matters, and level 4 develops and maintains communication on complex matters in complex situations. A manager can compare that to a real conversation their employee had last week. "Excellent communicator" gives them nothing to compare.

Step 4: Validate With the People Doing the Job

Before publishing, test the wording against three or four real people already doing the role at different levels. If someone who should obviously be a level 4 doesn't recognise themselves in that description, the wording is wrong — not the person. Our library of worked framework examples is a reasonable place to borrow anchor phrasing from if you're starting with a blank page.

Step 5: Pilot Before Rolling Out

Run one team through a complete assessment cycle first. The anchors that cause managers to hesitate, argue or ask "does this count?" are the ones to rewrite — finding that out with thirty people is cheap; finding it out after a company-wide rollout is not.

Where Consistent Scoring Breaks Down

Central Tendency and Leniency Bias

Even a well-written scale gets undermined by predictable rater errors. Dartmouth's HR guidance defines central tendency bias as "the tendency to evaluate every person as average regardless of differences in performance", and leniency bias as "the tendency to evaluate all people as outstanding and to give inflated ratings rather than true assessments of performance." The same source names several others — strictness, the halo effect, contrast effect, first impression error, similar-to-me bias. None of these are solved by better wording alone; they're solved by someone other than the rating manager seeing the evidence.

The Forced Distribution Trap

The instinctive fix is to force a distribution — a fixed percentage per band, regardless of evidence. General Electric's "vitality curve" under Jack Welch is the best-known version: roughly the top 20% as top performers, 70% as vital, the bottom 10% managed out. Most employers that copied it have since walked away: Microsoft ended seven years of stack ranking on 12 November 2013, Ford dropped its version in 2001 after a discrimination settlement, and Adobe, Accenture and Goldman Sachs abandoned forced ranking between 2012 and 2016.

Three sourced numbers: SFIA's seven levels of responsibility, the NHS KSF's four levels per dimension, and 2013 as the year Microsoft dropped stack ranking

Forcing an outcome doesn't fix a measurement problem — it just moves the argument from "what score is this?" to "who has to fill the bottom band?", which is worse to have. The fix that holds up, covered below, is calibration: comparing evidence against the same written anchors, not imposing a curve on unreliable ratings.

One Scale Applied to Every Job Family

O*NET, the US Department of Labor's occupational database, keeps two separate scales for every skill: an Importance scale (1–5), measuring whether a skill matters to an occupation at all, and a Level scale (0–7), measuring how much of it is needed. Its own example: lawyers and paralegals both need speaking skill at equally high importance, but a lawyer needs a materially higher level for courtroom argument. Few competency frameworks make this distinction — they apply the same 1–5 scale to "SQL" and "collaboration" as if proficiency meant the same thing for a technical skill and a contextual behaviour. It rarely does.

No Calibration Step at All

The most common failure isn't a badly written scale — it's a reasonably good scale with no process for checking that different managers apply it the same way. That's the gap the next section closes.

A Calibration Process That Actually Works

Comparison of scoring with no calibration, where ratings drift apart between teams, against scoring with a calibration session, where managers compare evidence against written anchors

A calibration session is a meeting, held before ratings are finalised, where managers compare specific cases against the written anchors and against each other — not a vague "let's discuss performance" chat. A workable version has four fixed parts:

  1. Who attends. Managers whose teams' ratings will be compared, plus a facilitator with no direct reports in the room and no rating to defend.
  2. What evidence is brought. The specific behaviour or output behind any rating above or below the mid-point — not just the number. No evidence, no discussion; it gets sent back.
  3. How disagreement is resolved. Against the written anchor, not seniority or force of personality. The wording gets read aloud and the group places the evidence against it together.
  4. How often it runs. Every cycle, before anything is communicated to employees — not as an audit afterwards, when the only options left are awkward and public.

The mechanism that makes this work is comparison, not correction. Nobody is told they're wrong in the abstract; they're shown the same anchor applied to a case from another team, and asked whether their own rating would survive sitting next to it — a far less confrontational conversation than "your ratings are too lenient" delivered after the fact.

How StaffCircle Helps Managers Score Consistently

One Scale, Defined Once, Applied Everywhere

Proficiency levels are defined once inside the skills and development framework and inherited automatically by every role using that competency, so a manager on one site scores against exactly the same anchor wording as a manager anywhere else — not a locally-drifted copy of it.

Evidence Attached to Every Rating, Not Just a Number

Every proficiency score can carry the observation, sign-off or example behind it, on the same record — which is what makes a calibration session possible at all. Our guide to building an audit-ready skills matrix covers the same evidence-first principle for a full team view.

Calibration Views Across Teams and Sites

Rating distributions can be compared across managers, teams and sites before a cycle closes, surfacing the manager whose ratings cluster suspiciously tight or high before it becomes a pattern nobody notices until an exit interview says otherwise.

Built to Reach Frontline and Deskless Managers

The organisations struggling hardest with consistent scoring often have the most managers per employee and the least time for a workshop — multi-site, shift-based, deskless teams. Mobile access and Microsoft Teams integration mean an assessment and its evidence can be captured at the point of work, not reconstructed from memory at quarter end.

Final Thoughts

A proficiency scale isn't the finished product of a competency framework — it's the instrument two people have to point at the same reading on. That takes a level count chosen for a reason, levels written as something a manager can observe, and calibration to catch what good wording alone can't. Most frameworks manage the first and stop.

If your framework already has levels but managers still land in different places on the same evidence, the fix usually isn't a rewrite — it's adding the anchor detail and calibration step that were missing the first time. Book a demo to see how StaffCircle turns a proficiency scale into ratings that hold up under comparison.

FAQ

What are proficiency levels in a competency framework?

Proficiency levels are the graded scale — typically 4 to 7 tiers — that a competency framework uses to describe how capable someone needs to be, or currently is, in a skill or behaviour. Each level should be written as an observable behaviour so an assessor can compare real evidence against it, not a vague adjective like "good" or "excellent".

What is the difference between a proficiency level and a skill level?

The terms are used interchangeably in practice, with no universal industry distinction. Some organisations use "skill level" for hard, teachable capabilities like software or a language, and "competency" or "behavioural level" for judgement-based capabilities like communication, because the two often need different numbers of levels and different anchor styles.

How many proficiency levels should a competency framework have?

There is no single correct number. SFIA uses seven levels because it grades autonomy and influence across an entire career. The NHS Knowledge and Skills Framework uses four per dimension because it's applied consistently across hundreds of thousands of staff, where simplicity protects consistency. Choose the count based on how the scale will be used, not by copying a template.

Should I use a 4-point or 5-point rating scale?

A 4-point scale removes the neutral midpoint, forcing a genuine decision and reducing central tendency bias — rating everyone as average. A 5-point scale gives finer differentiation but needs stronger anchors and more calibration effort to stop the midpoint becoming a default. Larger, more distributed manager populations generally get more consistent results from fewer levels.

What is a Behaviorally Anchored Rating Scale (BARS)?

A Behaviorally Anchored Rating Scale defines each point on a scale with a specific, observable example of behaviour, instead of a generic adjective. It's the approach behind well-built scales like the NHS Knowledge and Skills Framework, and the single change most likely to improve consistency between managers scoring the same competency.

What is the Dreyfus model of skill acquisition, and does it apply to performance ratings?

Developed by Stuart and Hubert Dreyfus in 1980, the model describes five or six stages an individual moves through as they develop mastery over years: novice, advanced beginner, competence, proficiency, expertise, and sometimes mastery. It describes long-term individual skill development, not a scale designed for a manager to score someone against annually, so citing it to justify a five-level rating scale is a common but imprecise use of the model.

How many levels does SFIA use?

SFIA, the Skills Framework for the Information Age, defines seven levels of responsibility: Follow, Assist, Apply, Enable, Ensure/Advise, Initiate/Influence, and Set Strategy/Inspire/Mobilise, each defined consistently across autonomy, influence, complexity and business skills.

How many levels does the NHS Knowledge and Skills Framework use?

Every dimension of the NHS Knowledge and Skills Framework, including its six core dimensions such as Communication, is scored on four levels, each with its own bulleted behavioural indicators describing what that level looks like in practice.

What is central tendency bias in performance ratings?

Central tendency bias is the tendency to rate every person as average, regardless of real differences in performance. It shows up as a cluster of ratings around the midpoint and is one of the main reasons organisations introduce calibration sessions before finalising ratings.

What is leniency bias?

Leniency bias is the tendency to rate everyone as outstanding and give inflated ratings rather than a true assessment. Its opposite, strictness bias, rates everyone at the low end and is overly critical regardless of actual performance.

Why did companies abandon forced ranking and stack ranking?

Forced ranking, popularised by General Electric's "vitality curve" under Jack Welch, imposes a fixed distribution regardless of the underlying evidence. Microsoft ended seven years of stack ranking in November 2013; Ford, Adobe, Accenture and Goldman Sachs dropped similar systems between 2001 and 2016 — largely because forcing an outcome doesn't fix a measurement problem, it just shifts the disagreement onto who fills the bottom band.

What is a performance calibration session?

A meeting held before ratings are finalised, where managers compare the evidence behind specific ratings against the same written anchors and against each other, facilitated by someone without a stake in any individual rating. Disagreement is resolved by returning to the anchor wording, not seniority or opinion.

Should the same proficiency scale apply to technical skills and behavioural competencies?

Not always. O*NET, the US Department of Labor's occupational database, uses separate Importance and Level scales because how much of a skill is needed can vary independently of whether it matters at all. Hard, teachable skills often support finer scales reliably; behavioural competencies are harder to separate past four or five levels, so one identical scale can produce misleading precision.

How often should proficiency levels be reassessed?

Match the cycle to how often the underlying evidence changes rather than a fixed calendar. An annual cycle aligned to a calibration session works for most behavioural competencies; technical skills tied to a changing tool, process or regulation should be reassessed whenever that change happens.


About the author

Mark Seemann is the CEO and Founder of StaffCircle, the AI performance management platform for mid-sized organisations. He writes about performance management, employee development and the practical use of AI in HR. Connect with Mark on .