In a recent conversation, the HR director of a large pharmaceutical company described a performance appraisal system they are developing in which peers will be involved alongside the manager. The reasoning is straightforward: managers may be more likely to notice highly visible employees – including more extroverted colleagues – while someone performing equally well in the background may be seen more clearly by the people who actually work alongside them. The expectation is that bringing in more perspectives will produce feedback that is ‘more objective, transparent and comprehensive’.
If one manager can be biased, ask more people. Logical. After all, only the manager can be biased. The eight colleagues we ask obviously cannot.
There is good reason to collect those additional perspectives. Peer and subordinate ratings can contain relevant performance information that is missing from a manager’s rating. In one meta-analysis, both sources provided significant incremental information about objective performance measures beyond that provided by other rating sources. [1]
Involving more people can broaden the information available. At the same time, every additional rating brings another human judgement into the system – with its own observations, interpretations and possible biases.
When the result feeds into a performance rating, a talent decision, promotion or pay, it matters whether we have produced better evidence or simply collected more ratings. Still, at least it shows that we are taking the matter seriously: we have produced an impressive amount of data.
In two large samples of managers receiving multisource developmental ratings, individual-rater effects accounted for 62% and 53% of the variance in ratings. [2] Later work challenged aspects of the model, while still finding substantial individual-rater effects. [3]
So adding peers does not simply add information. It adds raters.
Peer ratings therefore do not, by themselves, remove the problems associated with human judgement. They may introduce additional rater effects without necessarily removing the problems in the manager’s judgement.
Then comes the question of what we do with the differences.
I will spare you the rest of the dry research findings, but the nerds will find them in the references.
Reading this back, it looks pretty much like a critique, so you would be justified in asking: what now? Should we stop asking other people? Should we ignore what colleagues who actually work with someone can see? How would I do it?
There is not enough space on this platform to address that question properly and with sufficient evidence. If you are interested in the answer, I discuss it in detail in Judging Strangers: The Readability Trap in Hiring – including how human observations can be turned into decision-relevant information without confusing more data with better evidence.
Sources
[1] Conway, J. M., Lombardo, K. & Sanders, K. C. (2001). A Meta-Analysis of Incremental Validity and Nomological Networks for Subordinate and Peer Rating. Human Performance, 14(4), 267–303. DOI: 10.1207/S15327043HUP1404_1.
[2] Scullen, S. E., Mount, M. K. & Goff, M. (2000). Understanding the Latent Structure of Job Performance Ratings. Journal of Applied Psychology, 85(6), 956–970. DOI: 10.1037/0021-9010.85.6.956.
[3] Hoffman, B. J., Lance, C. E., Bynum, B. & Gentry, W. A. (2010). Rater Source Effects Are Alive and Well After All. Personnel Psychology, 63, 119–151. DOI: 10.1111/j.1744-6570.2009.01164.x.
[4] Conway, J. M. & Huffcutt, A. I. (1997). Psychometric Properties of Multisource Performance Ratings: A Meta-Analysis of Subordinate, Supervisor, Peer, and Self-Ratings. Human Performance, 10(4), 331–360. DOI: 10.1207/S15327043HUP1004_2.
[5] Harris, M. M. & Schaubroeck, J. (1988). A Meta-Analysis of Self-Supervisor, Self-Peer, and Peer-Supervisor Ratings. Personnel Psychology, 41(1), 43–62. DOI: 10.1111/j.1744-6570.1988.tb00631.x.
[6] Mount, M. K., Judge, T. A., Scullen, S. E., Sytsma, M. R. & Hezlett, S. A. (1998). Trait, Rater and Level Effects in 360-Degree Performance Ratings. Personnel Psychology, 51(3), 557–576. DOI: 10.1111/j.1744-6570.1998.tb00251.x.
[7] Batista-Foguet, J. M., Saris, W., Boyatzis, R. E., Serlavós, R. & Velasco Moreno, F. (2019). Multisource Assessment for Development Purposes: Revisiting the Methodology of Data Analysis. Frontiers in Psychology, 9, 2646. DOI: 10.3389/fpsyg.2018.02646.



