How Performance Reviews Reinforce Existing Bias

Why the Room Gets Quieter for Some Voices

Farah Adebayo has spent over a decade watching performance reviews shape careers, and the pattern is impossible to ignore. The same people who receive glowing feedback year after year often share a quiet commonality: they fit a narrow, unspoken mold of what “professionalism” looks like. Meanwhile, employees who code-switch, who navigate dual cultural identities, or who simply communicate differently find their reviews littered with vague critiques—”needs to be more strategic,” “not assertive enough,” “too direct.” These aren’t assessments of output. They are judgments of identity, refracted through a system that was never designed to be neutral.

Performance reviews promise objectivity. They come with rating scales, competency frameworks, and calibration sessions meant to strip away subjectivity. Yet a growing body of research reveals what many workers have long sensed: the process actively preserves the biases it claims to eliminate. When a manager sits down to evaluate a direct report, they are not measuring performance in a vacuum. They are measuring their own comfort, their own assumptions about leadership potential, and their own interpretation of behaviors that are deeply shaped by race, gender, and cultural background.

A diverse team seated around a conference table, one person speaking while others listen with varied expressions

The Research That Maps the Broken Feedback Loop

Academic studies have documented a consistent, troubling gap. A landmark analysis of performance evaluations in the technology sector found that women and people of color receive significantly less actionable feedback than white men. When they do receive feedback, it skews heavily toward personality traits rather than work outcomes. A Black woman might be told she needs to “soften her tone,” while a white male peer receives concrete guidance on project management. The critique is not about the work; it is about her presence in the room.

This is not a matter of a few biased managers slipping through the cracks. The structure of most review cycles intensifies cognitive shortcuts. Managers are asked to recall months of work, summarize it into a single score, and justify that score in writing—often under tight deadlines. In that scramble, the brain leans on pattern recognition. It defaults to the prototype of a “high performer” that culture has reinforced: someone who speaks up in meetings, who self-promotes comfortably, who mirrors the communication style of those already in power. Anyone who deviates from that prototype faces a steeper climb, not because of their results, but because they don’t trigger the same automatic confidence.

When “Culture Fit” Becomes a Gatekeeper

Many organizations embed questions about “alignment with company values” into their review forms. On the surface, this seems reasonable. In practice, it becomes a powerful vector for homogeneity. What does “living the value of collaboration” look like? To one evaluator, it might mean volunteering for high-visibility projects. To another, it might mean disagreeing openly in a brainstorm. But without rigorous, behavioral anchors, the rating defaults to the evaluator’s personal comfort zone. An introverted engineer who builds quiet consensus through one-on-one conversations might be marked down for lacking “executive presence,” while an extroverted marketer who dominates meetings earns top marks for “leadership.”

This dynamic hits hardest at the intersections. Farah Adebayo’s research-informed perspective highlights how Black women, for instance, face a double bind. Assertive behavior that is praised in male leaders gets coded as “angry” or “difficult” when exhibited by Black women. Collaborative behavior, on the other hand, can be read as a lack of authority. There is no safe middle ground, because the scoring criteria were built for a different person entirely. The review form itself becomes a mirror that reflects back not the employee’s contributions, but the evaluator’s internalized stereotypes.

A woman of color sits at a desk reviewing a document, her expression thoughtful and measured

The Myth of the Calibration Session

Calibration meetings are often held up as the fix for inconsistent ratings. Managers gather, share their proposed scores, and challenge each other’s reasoning. The intent is to balance out individual bias. The reality is more complicated. Group dynamics simply shift the bias from the individual to the collective. If most managers in the room share similar backgrounds and unexamined assumptions, calibration reinforces rather than corrects. A study in the Journal of Applied Psychology found that calibration discussions can actually amplify gender and racial disparities when the group lacks diverse representation among the calibrators themselves. The loudest voices in the room, often those of senior leaders, set the standard for what “strong performance” sounds like, and everyone else adjusts their ratings toward that norm.

Even the language used in these meetings reveals the problem. Descriptors like “polished,” “commanding,” and “gravitas” get attached to candidates who look and sound like the evaluators. Employees who deliver equally strong results but present differently receive more tentative endorsements: “solid contributor,” “needs more time to develop,” “good cultural steward.” These linguistic patterns are not random. They are the fingerprint of systemic bias, pressed into every performance cycle.

The Documented Cost of Ambiguous Criteria

When performance criteria are vague—think “demonstrates innovation” or “acts as a team player”—evaluators fill in the gaps with their own prototypes of what those behaviors look like. Research from the Center for WorkLife Law identifies this as “prove-it-again” bias: groups that are stereotyped as less competent must provide more evidence of their skills to receive the same rating. A white man might be presumed innovative on the basis of one successful pitch; a woman of color might need to deliver three measurable breakthroughs just to be seen as “meeting expectations.” Over time, this discrepancy compounds into promotions, pay gaps, and attrition. The review that feels like a minor slight to the manager becomes a career-altering tax for the employee.

Farah Adebayo has seen this pattern survive multiple “redesigns” of performance management. Companies shift from annual reviews to quarterly check-ins, from numerical ratings to qualitative feedback, from manager-only evaluations to 360-degree feedback. Yet the bias persists because the conversation never fully addresses the core issue: who defines what good performance looks like, and whose standards are being applied. Without interrogating the origin of the criteria, new formats simply repackage old prejudices in cleaner templates.

A close-up of hands writing notes on a performance review document, a pen resting on the paper

Reimagining the Review Without Abandoning Accountability

The answer is not to scrap performance reviews entirely—though many burned-out HR leaders have considered it. Feedback, when delivered well, drives growth. The challenge is to strip bias from the architecture of evaluation. That starts with behavioral specificity. Instead of rating “communication skills,” a review should assess whether an employee “summarized technical concepts for non-technical stakeholders in three major projects.” Instead of measuring “leadership,” it should capture whether the person “mentored two junior colleagues through their first client presentations.” Vague traits become observable actions, and observable actions are harder to distort with stereotype.

Second, organizations must audit their review data with the same rigor they apply to financial reports. If aggregated scores show consistent disparities by race, gender, or tenure, that is not a pipeline problem—it is a process problem. A commitment to equity requires analyzing the language in written feedback, tracking who receives developmental versus personality-based critiques, and holding calibrators accountable for the patterns they produce. Transparency in this audit is non-negotiable. When managers know their own rating histories will be reviewed for bias, the incentive to default to prototypes weakens.

Training Evaluators to Recognize Their Own Filters

Most bias training fails because it treats unconscious bias as a personal flaw rather than a predictable cognitive response. Effective interventions help managers recognize when they are substituting a vague impression for a concrete observation. They teach evaluators to ask themselves a simple question before finalizing any rating: “What specific, observable behaviors am I basing this on?” If the answer trails off into generalities about “presence” or “style,” the evaluation needs revision. This practice, combined with structured note-taking throughout the year rather than rushed end-of-cycle recollection, reduces the memory gaps that bias exploits.

Organizations also need to widen the lens on who gets to evaluate. Single-rater systems magnify individual bias; multi-rater systems distribute it. When feedback comes from peers, direct reports, and cross-functional partners, no single evaluator’s prototype dominates. The aggregation process itself can surface outliers—a manager who consistently rates women lower than men on the same competency—and prompt a deeper conversation. This is not about diluting accountability. It is about recognizing that one person’s impression is not the same as objective truth.

What a Bias-Aware Review Actually Feels Like

An employee who has experienced a fair review often describes it not as “nice” but as recognizable. They see their actual work reflected back to them, not a caricature filtered through someone else’s expectations. Their feedback mentions the late-night troubleshooting, the client relationship they repaired, the process improvement they quietly implemented. It does not comment on how “confident” they appeared or whether they “smiled enough” in the last all-hands. That specificity is the difference between an evaluation that develops and one that diminishes.

For Farah Adebayo, the goal is not perfection but a system that corrects itself. No process will eliminate every instance of bias. But a process can be designed to catch its own errors, to make the invisible visible, and to stop treating a manager’s comfort as a valid performance metric. The evidence is clear that current reviews reinforce existing hierarchies. The question is whether organizations are willing to rebuild the machinery, not just repaint it.

Frequently Asked Questions

Why do performance reviews often reflect bias instead of actual work?

Performance reviews rely on human judgment, which is shaped by cognitive shortcuts and cultural prototypes. When criteria are vague—like “demonstrates leadership”—evaluators unconsciously fill in the gaps with their own image of an ideal employee. That image is often influenced by race, gender, and communication style, so behaviors that differ from the prototype get rated lower, regardless of results.

Can calibration meetings fix biased performance ratings?

Calibration meetings can help if they are structured with diverse evaluators and clear behavioral anchors. Without those safeguards, they risk amplifying bias by aligning ratings with the dominant group’s perspective. Research shows that when calibrators share similar backgrounds, they tend to reinforce each other’s assumptions rather than challenge them.

What is the most effective way to reduce bias in performance reviews?

Shifting from trait-based ratings (“communication skills”) to behavior-based evidence (“led three client presentations that retained at-risk accounts”) is the single most impactful change. Pair this with multiple raters, year-round note-taking by managers, and regular audits of rating patterns by demographic group to create a system that self-corrects.

Does bias show up differently for people with multiple marginalized identities?

Yes. Intersectional bias means that a Black woman, for example, may face a double bind where assertive behavior is penalized as “angry” while collaborative behavior is seen as lacking authority. The same actions that earn praise for one group can become a liability for another, because the evaluation criteria are not neutral—they are rooted in the dominant culture’s expectations.