Performance Management

What Makes the Best Mac and Cheese?

Why performance ratings need shared criteria, like judging a mac and cheese cook-off, and how to calibrate managers and employees on the same scale.
Sarah Katherine Schmidt
VP of Customer Experience

Ask five people what makes the best mac and cheese and you'll get five different answers. One person wants it baked with a crunchy top. Another wants it stovetop and molten. Someone will insist real cheddar is non-negotiable, someone else swears by a cheese blend, and at least one person is quietly a Velveeta loyalist and unwilling to admit it in public.

None of those opinions are wrong. But if you're running a cook-off and you haven't agreed on what you're actually judging, "best" doesn't mean anything. Creaminess? Cheese pull? Crunch-to-sauce ratio? Until the judges agree on criteria, a 9 from one judge and a 9 from another could be describing two completely different dishes.

This is exactly the problem most organizations run into with performance ratings.

The Professor Problem

Think back to college. Some professors graded conservatively, some handed out A's like candy, and a "B+" meant something different in every classroom. Nobody coordinated on what the grade actually represented, so the grade itself became unreliable. You could work just as hard in two different classes and walk away with two very different letters.

Performance ratings have the same failure mode. Without a shared standard, "Strong Performer" from one manager might describe what another manager would call "Developing," and an employee's rating ends up depending more on who their manager is than on what they actually did. That's not fair to employees, and it quietly erodes trust in the whole review process.

The fix is what's often called norming: getting everyone, managers and employees alike, calibrated on what each rating actually means before anyone starts filling out a form. It's the equivalent of the cook-off judges sitting down beforehand and agreeing that they're scoring creaminess, cheese pull, and crunch, in that order, before a single bite is taken.

A Shared Scale, Defined in Plain Language

A typical rating scale runs from Inconsistent Performer up through Developing, Strong, and Outstanding Performer. The labels are simple, but the value is in defining, out loud, what separates one from the next.

RatingWhat it looks likeCalibrating questionsInconsistent PerformerRarely meets the expected bar for the role, and some job functions may need real development or improvement. For a newer employee, this often just means they're still learning the role, not that something is wrong.Is this person effective in some parts of the job but not others? Do they have a skill gap relative to what the role requires?Developing PerformerShows many of the right behaviors but still needs growth to consistently hit the bar. Everyone is developing something, regardless of level, so this rating isn't a red flag on its own. Worth remembering: consistent doesn't automatically mean strong.Are there workstreams this person is still learning? Are they producing consistently high-caliber work yet, or getting there?Strong PerformerConsistently meets the high bar for the role.Does this person reliably take ownership of complex, interdependent work from start to finish? Would you recommend them to a peer team without hesitation?Outstanding PerformerConsistently exceeds the bar and serves as a model for others in the organization. This rating is meant to be rare, and it doesn't automatically mean someone is ready for a promotion. It might just mean they're ready for more varied, challenging work.Is this someone others in the organization already look to as a leader? Could you pair them with someone who's struggling and expect real improvement as a result?

The Same Questions, Asked From Both Sides of the Table

Here's the part that makes norming actually work: the questions above aren't just for managers. When employees are asked to reflect on their own performance using the same criteria, self-assessment and manager assessment start speaking the same language. An employee asking themselves "am I still learning the skills required to operate independently in this role" is running the same check their manager is running. That overlap is what keeps ratings from feeling arbitrary or one-sided. It's also why norming sessions tend to work best as shared exercises, sometimes literally, with breakout groups practicing on sample scenarios before anyone applies the scale to a real review.

What the Ratings Look Like in Practice

Definitions only go so far. A few composite examples, built from real review patterns, make the calibration concrete:

Thomas is an Inconsistent Performer. He's adjusting to working across teams but is still struggling with some of the more technical parts of the role, particularly the discipline of shipping work in structured sprints. There's a real path to a higher rating here if he closes that specific gap.

Sally is a Developing Performer. She grew noticeably this year in communication and execution, and stepped into a role that required working across stakeholders at every level. Where alignment wasn't clear, though, she sometimes avoided the harder conversation instead of pushing through it. The growth edge for her is leaning into that friction instead of around it.

Priya is a Strong Performer. She picked up a new technology platform quickly, became the team's go-to expert on it, and then trained the rest of the team, which measurably improved team output. Colleagues seek her out for her expertise, and she keeps looking for ways to grow her own skills.

Domonick is an Outstanding Performer. He models the organization's values, takes initiative on the hardest projects, and operates with minimal oversight while still producing high-quality work. Just as important, he actively coaches and advocates for colleagues rather than just carrying his own load.

Notice that none of these examples turn on a single number or metric. They turn on a shared vocabulary, the same one the manager and the employee were both calibrated on.

Watching for the Signals Behind the Rating

Ratings are useful shorthand, but the specific behaviors underneath them matter more, especially for spotting retention risk early. Missed deadlines, unexpected skill gaps, and quality issues tend to cluster around Inconsistent ratings. Seeking out development opportunities and improving work quality without yet being fully independent tend to cluster around Developing. High productivity paired with real self-awareness and openness to feedback tend to cluster around Strong. And a diverse skill set combined with initiative on critical projects tends to cluster around Outstanding. None of these signals are the only indicator of performance, and what counts as a signal will vary by team and role, but naming them gives managers something concrete to watch for between review cycles rather than reconstructing an impression from memory once a year.

The Rating Isn't the Finish Line

The real point of calibrating ratings isn't to produce a clean label at the end of a review cycle. It's to create a shared, clear starting point for what happens next. That's what an individual development plan is for: not a one-time form, but an ongoing tool employees and managers build together to close specific gaps and set a longer-term career direction. The best version of this splits into two related pieces, a short-term growth plan tracked through concrete action items, and a longer-term career vision that gives the growth plan somewhere to point. Together, they turn a single rating into an actual development process instead of a verdict.

Back to the Mac and Cheese

Going back to that cook-off: the goal was never to get everyone to agree on one "best" recipe. It was to get the judges scoring the same things, so a 9 means the same thing across the table. Performance norming works the same way. Nobody expects every manager to see every employee identically, and they shouldn't. What norming buys you is a shared standard, so that a rating reflects the work someone actually did, not which manager happened to be grading that year.

No items found.
Performance Management
The Skill Gap Behind Every Avoided Conversation
Read More
HRTech
Why Most HRTech Rollouts Fail (And What Separates the Ones That Don't)
Read More
Performance Management
What performance management should look like at 25-250 employees, and the mistakes companies make at each stage.
Read More
Ready to turn ā€œwe should fix performanceā€ into a repeatable SMB growth system?
Starting at $3,588 / yr

A platform that drives
real results.

Get started