The Hidden Fault Lines: How Performance Reviews Reinforce Existing Bias

On paper, the performance review is a ritual of fairness. A manager sits with an employee, rubric in hand, to talk achievements, growth areas, and goals. The forms are standardized; the rating scales look objective. But if you’ve sat through enough of these, you know the unease in the room. The numbers don’t quite add up. The feedback feels oddly personal, anchored to impressions formed in the first five minutes of a meeting months ago. This isn’t a glitch. It’s what happens when human cognition meets a process that was never really built to be neutral.

Diverse team of professionals engaged in a collaborative meeting in a modern office

Farah Adebayo here. A big chunk of my career has gone into studying organizational behavior and how our minds construct reality. In most companies, the performance review isn’t a measurement tool. It’s a mirror reflecting the biases we carry into work. It doesn’t just record performance; it actively shapes it, often penalizing people who don’t fit a narrow, unspoken mold. Research from psychology, sociology, and management studies paints a clear picture: the process has structural flaws that perpetuate inequality. Until we acknowledge those fault lines, we’re just grading on a curve of our own making.

The Architecture of Subjective Judgment

To see why reviews fail, look at the cognitive shortcuts our brains grab. One of the most pervasive is the idiosyncratic rater effect. Researchers have shown that over 60% of a performance rating reflects the rater’s own characteristics, not the employee’s actual output. Your review says more about your manager’s personality, their leniency scale, and their personal definition of “leadership” than it does about your work. A 2015 study in Personnel Psychology found that a manager’s own performance history, their tendency to be agreeable or critical, even their mood on evaluation day, significantly skewed scores.

This subjectivity feeds confirmatory bias. Once a manager forms an early impression—maybe during the interview or a project kickoff—they subconsciously hunt for evidence that confirms it. If you got the “high potential” label early, your successes get magnified; your failures get blamed on external factors. If you were marked “average,” your wins become luck, and your missteps proof of your limits. This isn’t a conscious plot. It’s the brain’s energy-saving mode, using heuristics that skip deliberate analysis. The result is a feedback loop where early labels become self-fulfilling prophecies.

The Tightrope of Likeability and Competence

This loop is most punishing where identity meets perception. Research has shown again and again that women, especially women of color, walk a tightrope between being seen as competent and being seen as likeable. A classic Harvard Business Review study found that in performance reviews, women are much more likely to get vague feedback about their “style” or “tone,” while men get concrete, actionable feedback tied to technical skills. Words like “abrasive,” “aggressive,” and “emotional” show up in women’s reviews far more than men’s, even when controlling for job role and performance metrics.

This isn’t just a quirk of language. It has material consequences. Feedback that criticizes personality instead of output gives no clear path to improvement. Telling someone to “be more confident” or “soften her approach” isn’t a development plan. It’s a cultural demand to assimilate to a narrow behavioral standard. I’ve seen this pattern in tech firms, law offices, and academic institutions. People who deviate from the expected norm—in communication style, assertiveness, even their non-work interests—get penalized in ways that are hard to quantify but impossible to ignore over a career.

Black female professional leading a discussion with colleagues around a glass conference table

Proximity Bias and the Myth of Visibility

The shift to hybrid and remote work has exposed another layer of distortion: proximity bias. Managers consistently rate employees they see more often as higher performers, even when objective metrics show no difference in output. A 2022 survey by the Society for Human Resource Management found that 67% of supervisors admitted to factoring in-office presence into promotion decisions, often subconsciously equating physical visibility with commitment and productivity.

This bias doesn’t only hit remote workers. It hits anyone whose schedule or work style doesn’t match the manager’s. Parents who leave at 5 p.m. sharp, employees who prefer early morning or late evening hours, people who eat lunch at their desks versus those who network in the cafeteria—all face a quiet calculus of “cultural fit” that has nothing to do with results. The performance review, with its emphasis on “demonstrating initiative” and “being a team player,” becomes a trap. Those phrases are often code for “I see you around.” They systematically disadvantage anyone whose life circumstances don’t allow for constant, informal facetime.

The Recency and Halo Effects

Human memory isn’t a video recorder. It’s a storyteller that prioritizes the emotional and the recent. The recency effect means a project finished in the last two weeks before a review will disproportionately shape the overall rating, while a major success from eight months ago fades. The halo effect makes one positive trait—maybe an employee is exceptionally articulate in meetings—cast a glow over the entire evaluation, masking weaknesses elsewhere.

On the flip side, the horns effect lets a single negative incident, like a missed deadline or a tense email, taint the whole review period. These cognitive distortions aren’t evenly spread. Research suggests the halo effect often benefits people who fit the stereotypical image of a leader in their industry—often white men in suits in corporate settings, or the “brogrammer” archetype in tech startups. Meanwhile, the horns effect can stick to people branded as “difficult” for advocating for themselves or their team. That label follows them from review to review, no matter their actual performance trajectory.

Exhausted employee resting head on desk in a dimly lit office, symbolizing review stress

Rating the System, Not the Person

If the individual review is this flawed, what’s the alternative? The answer isn’t simply training managers to be less biased. Decades of data show that unconscious bias training, on its own, has a negligible effect on behavior. Sometimes it backfires by making people think they’re now immune to bias. A more effective shift means changing the evaluation’s structure.

Some organizations are moving toward real-time feedback systems that separate performance conversations from compensation and promotion decisions. By gathering short, specific input from peers, direct reports, and supervisors weekly or monthly, the recency effect loses its power. The feedback becomes about the work, not the person’s perceived trajectory. Companies like Adobe and Deloitte have famously ditched annual ratings for frequent “check-ins,” and early data suggests these models can reduce turnover among underrepresented groups.

Another approach redesigns rating scales to measure specific, observable behaviors instead of vague competencies. Instead of asking a manager to rate “leadership”—a loaded term—the rubric might ask: “How often does this person share credit with team members?” or “Does this person provide clear, written agendas before meetings?” This shifts the conversation from subjective judgment to factual observation, leaving less room for the idiosyncratic rater effect to operate.

The Uncomfortable Question of Purpose

Any attempt to fix performance reviews eventually has to sit with a deeper question: what is the review actually for? If the goal is allocating bonuses and identifying the bottom 10% for forced ranking, the process is inherently adversarial. It incentivizes managers to justify their pre-existing beliefs and employees to game the system. If the goal is genuine development, the review must be separated from compensation and conducted by someone with no stake in the budget.

My research-informed take is that we often expect a single tool to do two contradictory things: be a fair judge and a supportive coach. The courtroom and the classroom need different architectures. Until organizations have the courage to split these functions, performance reviews will remain a theater of equity, performing objectivity while quietly reinforcing the same old hierarchies. The data is solid. The lived experiences are consistent. The only question is whether we’re ready to build a system that matches the complexity of human performance, or whether we’ll keep pretending a form and a number can capture it.

Frequently Asked Questions

Why do performance reviews often feel more personal than professional?

Because they’re filtered through the manager’s cognitive biases. The idiosyncratic rater effect means the review reflects the manager’s personality and mental shortcuts as much as your actual work. Vague feedback about “style” or “attitude” often comes from subconscious expectations about how certain people should behave—having little to do with objective performance metrics.

Can’t bias training fix these problems?

Awareness is a starting point, but research shows standalone bias training rarely changes long-term behavior. The brain’s shortcuts are too ingrained. Effective solutions need structural changes to the review process—like collecting real-time, behavior-specific feedback from multiple sources and separating development conversations from salary decisions.

How does remote work make bias in reviews worse?

Remote and hybrid setups amplify proximity bias, where managers unconsciously rate in-person employees more favorably. This extends to anyone whose schedule differs from the manager’s—parents, caregivers, people with non-traditional hours. Visibility gets mistaken for commitment, and informal facetime becomes an unspoken metric in evaluations.

What is the most effective alternative to annual reviews?

Systems that emphasize frequent, future-focused check-ins centered on specific tasks and behaviors, not personality. When feedback is gathered continuously from multiple viewpoints and tied to observable actions, it reduces the impact of recency and halo effects. The key is to use these conversations for growth, not as a justification for a pre-decided rating or ranking.