Calibrating your interviewers: the panel is an instrument
You would not run a credit model you never validated. Most panels are exactly that.
Organizations audit every instrument they use except the interview panel. Yet interviewers vary enormously: some are harsh, some generous, some swayed by sector pedigree, some by confidence. Uncalibrated, the panel's output is noise around each interviewer's personal baseline.
Measure your interviewers
If your ATS stores scores, analyze them: which interviewers are systematically hard or easy, whose scores correlate with eventual performance, who never uses the top of the scale. Even without data, the debrief reveals them: the person whose "strong yes" has never been wrong, the one who has never given one.
Calibrate with a common candidate
- Have interviewers co-interview periodically and compare scores on the same candidate
- Discuss the divergences: what evidence did each person weigh?
- Anchor the scale: agree what a 3, 4 and 5 look like with real examples
Weight by track record
Some interviewers are genuinely better predictors - usually those who score on evidence and revise on data. Give their assessments more weight in debriefs, explicitly. This feels undemocratic; it is simply measurement. A panel is an instrument panel, and instruments get calibrated or they mislead.