Est.

Reference Check Processes That Predict Job Performance

Senior Contributor · · 9 min read
Cover illustration for “Reference Check Processes That Predict Job Performance”
Assessment Science · October 7, 2026 · 9 min read · 2,016 words

The low validity scores attached to reference checks come from one specific body of research. That research studied a specific kind of reference check, not the method itself. That research compared unstructured reference checks against structured interviews and work samples, and it placed the unstructured version meaningfully behind both as a predictor of job performance. The finding was accurate. What got lost in translation was the qualifier: most practitioners absorbed "reference checks don't work" and dropped the word "unstructured" somewhere along the way. Reference checks became a compliance ritual performed to limit legal exposure, and one well-known criticism in the hiring world holds that reference checks are performative. That criticism is fair when it's aimed at the unstructured version of the practice. When a reference check is run as a box to check before an offer goes out, it functions exactly like a box to check. The criticism describes how the tool gets used, not what the tool is capable of, and the distinction between structured and unstructured reference checking is the one that determines which description applies.

What changed when researchers studied structured reference checks instead

When researchers turned to structured reference checks in the 21st century, what they found changed the validity picture substantially. Zimmerman and colleagues looked in 2010, Hedricks and colleagues in 2013, and Taylor and colleagues in 2004, and each found strong correlations between structured reference checks and actual work performance, a sharp departure from the earlier verdict on unstructured checks. Those findings put structured reference checking in the same company as other well-regarded noncognitive predictors used in hiring, including personality tests, assessment centers, and biodata, all of which carry similar standing in terms of criterion-related validity. The U.S. Office of Personnel Management's assessment guidance documents that reference checks, done well, predict supervisory ratings of job performance, success in training, promotion potential, and employee turnover, and it notes that adding structure improves reference checks the same way it improves employment interviews.

Structure adds something else, too: incremental validity on top of tools already common in hiring systems, including cognitive ability tests and self-report personality measures. A structured reference check is not redundant with those tools. It captures information they cannot reach on their own. Have a supervisor or colleague rate a candidate's personality, and that observer rating will consistently beat the candidate's own self-reported personality at predicting how they perform on the job. A candidate describing their own work style in an interview cannot substitute for that outside view, no matter how self-aware the candidate is.

What "structure" in a reference check requires

Structure is a cluster of design decisions that, together, strip out the variability that makes unstructured checks weak predictors. The Office of Personnel Management identifies three core strategies: base the questions on a job analysis, ask every candidate's referees the same set of questions, and give the people conducting the checks standardized procedures for collecting and rating the data they get back. Each piece solves a different problem. Job analysis ties the questions to what the role actually demands. If you ask every referee the same questions, the conversation can't drift toward whatever feels comfortable to discuss. If rating procedures are standardized, a checker's gut feeling can't override what the data actually says.

Canada's Public Service Commission lays out a structured reference checking guide that breaks the process into seven stages: initiating the discussion, asking preliminary and verification questions, asking competency-based questions, assessing overall suitability, gathering additional comments from the referee, closing the interview, and recording additional comments from the reference checker afterward. That staging matters because it shows structure governing the whole arc of the conversation, not just the wording of individual questions. Compare that to the familiar unstructured approach, the one built around a single open-ended prompt like "tell me what it was like to work with this person." That question invites a referee to share whatever comes to mind, favorable or not, with no guarantee that two referees asked about two different candidates are even talking about comparable things. A structured process replaces that improvisation with a Reference Checking Form, which the guide recommends building to suit any occupational level and the needs of the specific hiring organization. The form covers verification questions, sections for assessing particular competencies, an overall suitability rating, and room for the checker's own observations. Getting the form right starts with getting the questions right, and that's the design choice that carries the most weight in the whole process.

How to design questions that generate predictive signal

A question generates predictive signal when it asks about a specific past behavior in a context that resembles what the target role actually demands. The job analysis sets the boundaries here: every question on the form should map onto a competency the role requires, so that every answer collected tells the hiring team something about a trait the job actually calls for. Canada's framework splits questions into two categories. Preliminary and verification questions just confirm the factual claims a candidate made on their application or resume. Competency-based questions ask the referee for behavioral evidence: what did this person do, in what circumstances, with what result?

Verification questions do useful, if narrow, work. They catch the gap between what a candidate claimed and what a former supervisor reports, which makes them a reasonable screen against resume inflation or outright fabrication. Competency questions carry the real predictive weight. Ask a referee to describe a time the candidate had to resolve a conflict with a difficult stakeholder, and you get a concrete account of behavior under a specific condition, the same kind of evidence that makes structured interviews outperform unstructured ones. Asking the same referee whether the candidate was "a good communicator" produces an adjective, not evidence.

The question set has to stay fixed for every candidate competing for a given role. Changing the questions from one candidate's reference check to the next introduces the same problem that plagues unstructured interviews: the hiring team ends up comparing answers to different prompts. If you pair a fixed question set with a standardized rating scale, the checker gets a real basis for combining responses across multiple referees, instead of falling back on a general impression of how the calls went.

Question design also has to account for what referees tend to volunteer on their own. Hedricks and colleagues found that referees bring up soft skills, like working well with others and communicating clearly, far more often than hard skills like computer programming or mathematical ability, and that which skills referees emphasize shifts depending on the job in question. Left to their own devices, referees gravitate toward interpersonal impressions. A question set built for a technical role has to ask directly about technical competencies, because referees won't volunteer that information on their own. Skipping that step quietly loses the data a technical hire most needs evaluated.

Why who you ask matters

The relationship between a referee and a candidate shapes how much useful signal that reference can produce, and it matters as much as the quality of the questions themselves. The Office of Personnel Management identifies supervisors, peers, and subordinates as the key referee categories, and each one observes a different slice of how a candidate actually works. Direct supervisors carry particular weight because they watched the candidate get measured against real accountabilities, not just interpersonal behavior in passing. A supervisor reference works more like a high-quality professional referral than a character vouch. Canada's guide tells hiring organizations to give candidates clear instructions when they're choosing referees: what kind of relationship counts, how recent the working relationship needs to be, and how relevant that person is to the specific competencies under review.

The most common objection to reference checking is that candidates only hand over referees who will say nice things, and that objection doesn't hold up against the evidence. Zimmerman and colleagues found in 2010 that even when referee pools were chosen entirely by the applicants, enough variance still existed across candidates for structured reference checks to predict work performance. Observer ratings of personality drawn from referees the applicant nominated still outperformed self-reported personality as predictors of job performance. Selection bias is a real feature of applicant-nominated referee pools, but it isn't fatal to the validity of what gets collected. The objection carries much more force against unstructured checks, where a glowing but vague reference gives you almost nothing to work with. A structured question set pulls differentiated, usable detail even out of a referee who likes the candidate and wants them to get the job.

Require supervisory references instead of settling for peer or personal ones, and you strengthen both the validity of the data collected and its resistance to uniform leniency. If a candidate can't or won't produce a former supervisor as a reference, that is handing the hiring team a signal before the process even starts. Treating supervisor contact as the default expectation, not an optional extra, builds that signal into the process from the outset. How a candidate assembles their list of referees is itself a piece of data worth tracking, separate from anything a referee eventually says.

What compliance rates reveal

Diagram: Compliance Rate as a Pre-Answer Signal. Visualizes: Visualize a two-part finding from Hedricks and colleagues (2010s): even after controlling for what referees actually said, the share of nominated referees who responded to the request…

Whether a candidate's nominated referees actually respond to the reference request predicts job performance and involuntary turnover on its own, independent of what those referees say once they're reached. Hedricks and colleagues found that even after controlling for the actual ratings referees gave on survey items, the share of referees who complied with the request was positively related to the hired candidate's performance and negatively related to involuntary turnover. A candidate whose nominated referees mostly decline to respond is flagging a problem before a single competency question gets scored.

Capturing that signal takes a deliberate choice in how you build the process. An organization has to track not just the content of what referees say but who answers the request and who doesn't, which calls for a logging system built into the process, not an afterthought bolted onto the question form. If a candidate generates a low compliance rate among nominated referees, you should treat that as elevated risk, even when the responses that do come back are positive ones. This finding changes what reference checking actually is: a multi-signal assessment that starts the moment a candidate is asked to name their referees, well before any question gets asked of anyone.

Employers take on legal exposure from running reference checks poorly. The Office of Personnel Management's guidance defines negligent hiring as a failure to exercise reasonable care when bringing on a new employee, and it notes that conducting reference checks can lower the risk of exactly that kind of lawsuit. One negligent hiring case in the research involved a service technician with a documented history of harassment and forgery, information that a proper reference check would have surfaced. Plaintiff attorneys in that case argued a routine check would have disqualified him before he was ever hired. Negligent hiring doctrine makes thorough reference checking a way of managing risk, not a bureaucratic step to clear before extending an offer.

The fear of defamation that leads many employers to say as little as possible when they're contacted as someone else's reference creates its own exposure on the other side of the transaction. An employer who gives nothing but dates of employment forces the hiring organization to make a decision without the information negligent hiring doctrine expects them to have sought out in good faith. Most U.S. states extend qualified privilege protection to employers who provide reference information. So if an employer answers honestly and without malice, they are generally shielded from defamation claims even if some detail later turns out to be wrong. Federal employers also get added protection under the Federal Tort Claims Act, and that tilts the legal calculus even further toward giving a candid reference. Risk here runs in both directions. A thin reference doesn't eliminate legal exposure, it just moves that exposure from the employer giving the reference to the organization that hired without the information it needed.

Sources

  1. Reference Checking
  2. Structured reference checks - Canada.ca
  3. Reference Checking Guide
  4. Factors affecting compliance with reference check requests - Hedricks - 2019 - International Journal of Selection and Assessment - Wiley Online Library

More in Assessment Science