What is the 9-box grid?
A practical guide from the 9Box team · Last updated 22 August 2026
The 9-box grid is a talent review tool. You place every person in a group onto a three-by-three grid — performance along one axis, potential along the other — and end up with a single picture of where the group's strength, risk and untapped capacity sit.
The grid is usually traced back to work McKinsey did with General Electric in the 1970s, where a similar nine-cell matrix was used to decide which business units deserved investment. The version HR teams use today applies the same logic to individuals. That attribution is widely repeated but thinly documented, so treat it as background rather than provenance — nothing about whether the grid is useful depends on where it came from.
In short: a 9-box grid is a three-by-three matrix used in talent reviews. One axis rates an employee's performance in their current role as low, medium or high; the other rates their potential to grow into a larger role on the same three-point scale. Placing everyone in a group on the grid produces nine boxes, each carrying a different recommended action — invest, stretch, coach, or address.
The two axes
Performance (horizontal) — what this person has actually delivered. Performance is backward-looking and evidence-based. It asks a narrow question: over the review period just ended, how did this person do against the expectations of the job they hold right now? Results, quality, reliability, how they worked with others. You should be able to point at something specific for every rating.
Potential (vertical) — how much larger or more complex a role this person could grow into. Potential is forward-looking and, unavoidably, a judgement. It asks a different question: if the organisation gave this person materially bigger or different work, could they do it, and do they want it? The useful signals are learning speed, how they behave when the problem is unfamiliar, whether their influence already reaches beyond their own remit, and — the one most often skipped — whether they actually want a bigger job.
Each axis has three levels: low, medium and high. That gives nine boxes.
The point people miss: the two axes are measured against different reference points. Performance is measured against the bar for the role someone already holds. Potential is measured against a role they don't hold yet. This is why a brilliant senior specialist can be rated high performance and low potential without it being a criticism — it means they are excellent at what they do and are not on a path to a substantially different job. Teams that treat "low potential" as an insult end up rating everyone high, and the grid stops telling them anything.
What the nine boxes mean
The nine boxes at a glance
| Box | Performance / potential | What you do about it |
|---|---|---|
| Star | High / high | Retain deliberately, reward, and name them in a succession plan. |
| Growth Employee | Medium / high | Stretch assignment and exposure to decision-makers. |
| Rough Diamond | Low / high | Diagnose the blocker before you conclude anything about the person. |
| High Performer | High / medium | Deepen the expertise, broaden the scope. |
| Core Player | Medium / medium | Engagement and incremental growth. Do not ignore this box. |
| Inconsistent Player | Low / medium | State the bar in writing, then coach against it. |
| Trusted Professional | High / low | Mentoring, standards-setting, onboarding. |
| Effective | Medium / low | Recognise the contribution and retain the person. |
| Underperformer | Low / low | A time-bound improvement plan, a different role, or an exit. |
Box names and their order match the grid 9Box draws, so the vocabulary in this guide is the vocabulary in the tool.
High potential
Rough Diamond — low performance, high potential
Real capability that has not converted into results yet. Something is in the way: the wrong role, a bad manager fit, an unclear brief, or too little time in the job to have delivered anything. The action is diagnostic, not disciplinary — find the blocker before you conclude anything about the person.
Growth Employee — medium performance, high potential
Delivering solidly and visibly capable of more. The constraint is usually exposure rather than ability. Stretch assignments, work that puts them in front of people who make decisions, and a manager who will actually let go of something.
Star — high performance, high potential
Delivering now and ready for more. These people have the shortest fuse on retention: they are the easiest to recruit away and the easiest to take for granted. The action is deliberate — retain, reward, and build succession plans that name them.
Medium potential
Inconsistent Player — low performance, medium potential
The capability is there in flashes but the delivery is uneven, and nobody can predict which version turns up. The action is clarity before coaching: state the bar explicitly, in writing, and manage against it. Ambiguity is what keeps people in this box.
Core Player — medium performance, medium potential
Dependable, steady, and usually the largest group. This box carries most of the organisation's actual output, and it is the box most likely to be ignored in the meeting because it generates no drama. The action is engagement and incremental growth, not benign neglect.
High Performer — high performance, medium potential
Consistently strong in the current role with moderate room to grow beyond it. Deepen their expertise and broaden their scope — more surface area at the same level is often more valuable to them, and to you, than a promotion neither of you wants.
Low potential
Underperformer — low performance, low potential
Not delivering, and no visible trajectory that changes it. This is the box everyone avoids discussing honestly, which is exactly why it needs a decision: a concrete, time-bound improvement plan, a different role, or an exit. Leaving it unresolved is a decision too, and usually the worst one.
Effective — medium performance, low potential
Doing the job properly. Not going to do a much bigger one. There is nothing wrong with this box — most organisations depend on it. Recognise the contribution and retain the person.
Trusted Professional — high performance, low potential
Deep expertise, applied excellently, in a role they are well matched to. Often your most senior individual contributors. The value they can add beyond their own output is transferring what they know, so mentoring, standards-setting and onboarding are the natural asks.
How the grid is used in a talent review
The grid is not a scoring exercise. Its output is not really the nine boxes — it is a shared, comparable view of a population that several managers now agree on.
A typical cycle looks like this. Managers place their own people first, individually. Those placements are then brought into a calibration session where the managers, their common leader and an HR partner work through the grid together, and — crucially — managers have to justify their placements to peers who manage comparable people. That is where the value is created. A manager who rates everyone highly gets challenged. A manager who is unusually harsh gets challenged. Two people in different teams doing similar work get compared directly, often for the first time. Placements move during the conversation, and they should.
What comes out is a set of decisions: who needs a development action and what it is, who is a retention risk, which senior roles have no credible successor, and who needs a difficult conversation that has been postponed.
In succession planning specifically, the grid is an input and not the plan. What it gives you is a shortlist: the people with the capability and the appetite to take on materially larger work. What it cannot give you is readiness. High potential means someone could grow into a bigger role; it says nothing about whether they could do a named job next quarter. Keep readiness as a separate judgement — ready now, ready in one to two years, ready in three or more — recorded against the role rather than against the person. A succession plan that consists of the top-right box and nothing else is not a plan, it is a list of the people most likely to leave.
The most useful output is often the empty space. A high-potential row with nobody in it, in a team of forty, is a succession problem that was invisible before you drew the grid. A senior role with no credible successor anywhere on the page is the single most actionable thing a talent review produces, and it is the thing most likely to be skipped because it is nobody's individual development plan.
Read the step-by-step guide to running the session, including who should be in the room and what to do with the output — or go straight to the talent review FAQ if you have one specific question.
Honest limitations
Anyone selling you the 9-box grid without this section is selling you something. The method has real, well-understood weaknesses.
Potential is the weak axis. There is no agreed definition of potential and no reliable way to measure it. In practice it often gets rated on proxies: confidence, visibility, articulacy, willingness to relocate or work long hours, or how much the person resembles the people doing the rating. Every one of those correlates with things that have nothing to do with capability.
Recency bias distorts performance. A review covering twelve months is usually rated on the last two. A great Q4 lifts someone a box; a bad project in the final month drops them one. Written evidence gathered through the year is the only real defence, and almost nobody does it.
The grid amplifies bias rather than filtering it. Structure creates an impression of objectivity, but a biased judgement placed in a nine-cell matrix is still a biased judgement — now with an official-looking position. Patterns worth actively checking before you accept a grid: how people returning from parental leave are rated, how part-time workers are rated, and whether "high potential" tracks demographics more closely than it tracks output.
Labels stick to people. The boxes are convenient shorthand, which is precisely the problem. Once someone has been described as a Core Player in a meeting, that description tends to follow them into the next cycle and into decisions they never see. Assume anything written on a grid will eventually be seen by, or repeated to, the person on it.
Forced distribution breaks it. Requiring a fixed percentage in each box turns a discussion into a rationing exercise, and managers respond by trading placements rather than assessing people. If a genuinely strong team produces a top-heavy grid, that is information, not an error to be corrected.
It is a snapshot, not a verdict. The grid describes one group at one moment, with the information available. People move between boxes, and a change of manager, role or brief moves them faster than anything else.
It does not measure several things people read into it. The grid says nothing about flight risk, about how critical a role is to the business, or about how soon someone would be ready for a promotion. "Star" does not mean "ready now". If you need those, track them separately — do not infer them from a box.
Used well, the 9-box grid is a structured excuse for a conversation that most organisations otherwise never have. Used badly, it is a way of writing down prejudices in a grid and calling the result data. The difference is entirely in how the session is run.
How to run it well
Every weakness above has a countermeasure, and none of them is complicated. They are just skipped.
Define both axes in writing before anyone is rated. Not during the meeting, and not in the manager's head. Write what low, medium and high mean on each axis, in your own organisation's language, and circulate it. If nine managers hold nine private definitions of high potential, the calibration session becomes a two-hour argument about vocabulary.
Rate the axes separately, and in that order. Do performance first and finish it, then start potential from scratch. The most common single error is letting a strong performance rating drag the potential rating up because it feels harsh not to. When that happens both axes say the same thing and the grid collapses into a ranked list you could have written without it.
Demand one piece of evidence per placement. "She led the migration and it landed two weeks early" is evidence. "He's very strong" is not. Requiring evidence in writing, before the meeting, catches a surprising proportion of gut-feel ratings before they reach the room.
Ask for evidence from the first half of the period. This is the only practical defence against recency bias that works in a live meeting. If every example a manager offers is from the last six weeks, you are not rating a year.
Compare people, not definitions. Arguing in the abstract about whether someone is medium or high goes nowhere. Naming two specific people and asking whether one is genuinely operating above the other resolves it in a fraction of the time.
Read the shape before you close. Step back and look at the distribution rather than the individuals. Is the top-right corner crowded — ratings inflation, most likely, rather than an exceptional team? Is the high-potential row empty? Does high potential track demographics, or volume, or which manager is most persuasive, more closely than it tracks output? Someone in the room has to be willing to ask that out loud, and it is usually the HR partner's job.
Attach an owner and a date to every placement. The grid is not the deliverable. A development action, a stretch assignment, a coaching conversation, a role change or a postponed difficult conversation — each with a named owner and a date — is the deliverable. Without that you have held a two-hour meeting and produced a picture.
Write it as though the person will read it. Because sooner or later someone will: a subject access request, a leaver's handover, a forwarded slide. Notes you would not stand behind in front of the person are notes that should not be on the grid.
The step-by-step guide turns each of these into a numbered step, with an agenda and timings for the session itself.
Frequently asked questions
What are the two axes on a 9-box grid?
Performance and potential. Performance is backward-looking: what the person delivered during the review period, measured against the expectations of the role they currently hold. Potential is forward-looking: how much larger or more complex a role they could grow into, and whether they want one. Each axis has three levels — low, medium and high — which produces nine boxes.
What is the difference between performance and potential?
They are measured against different reference points, and that is the distinction most teams get wrong. Performance is measured against the bar for the job someone already has. Potential is measured against a job they do not have yet. This is why an excellent senior specialist can be rated high performance and low potential without it being a criticism — it means they are very good at what they do and are not on a path to a substantially different role.
How often should you run a 9-box review?
Most organisations run one annually, often alongside the performance cycle, and some add a lighter mid-year check. More frequently than twice a year and the placements stop moving enough to be worth the meeting. Less frequently than annually and the grid goes stale, because people change roles, managers and circumstances faster than that.
Should you tell employees which box they are in?
There is no single right answer, but quoting box names to people is rarely the useful option. The labels are shorthand designed for a calibration discussion, not feedback, and they tend to become permanent descriptions of a person rather than a judgement made at one moment. Most organisations share the development actions and the reasoning behind them without naming a box. What matters is deciding your approach explicitly and applying it consistently, because inconsistency is how the labels leak anyway.
Should we force a distribution across the nine boxes?
No. Quotas turn assessment into rationing, and managers respond by trading placements rather than assessing people. If a genuinely strong team produces a top-heavy grid, that is information worth having. The exception is worth noting: if every team produces a top-heavy grid, the problem is your definitions, not your distribution, and the fix is to rewrite them rather than to impose a quota.
What are the main criticisms of the 9-box grid?
The potential axis has no agreed definition and is often rated on proxies like confidence, visibility or availability rather than capability. Recency bias means a twelve-month period gets rated on its last two months. The structure looks objective, so it can lend authority to a biased judgement rather than filtering it out. And the labels stick to people long after the meeting that produced them. None of these are reasons to avoid the method, but a review run without accounting for them produces confident-looking conclusions that are not reliable.
More questions — how big a grid should be, how long the session takes, whether the grid can drive promotion decisions — are answered in the 9-box grid and talent review FAQ.
Build one without the spreadsheet
9Box gives you all nine boxes pre-named, drag-and-drop placement, and a PNG or PDF export for the meeting. $49 a month or $490 a year — and while we finish building billing, signing up gives you the whole thing.
Build your first grid