Cognitive Ability Tests and Adverse Impact in Hiring
Tests that predict job performance best also create the largest racial score gaps.

You can't fix with good intentions the bind that cognitive ability tests put HR practitioners in. The traits that make these tests some of the strongest predictors of job performance are the same traits tied to the score differences between racioethnic groups that trigger adverse impact. Campion & Campion (2025) state that the hiring procedures that are most valid and most affordable, meaning mental ability tests, tend to produce the largest score gaps between racioethnic subgroups, and that gap creates real legal exposure under anti-discrimination law. The mechanism behind this depends not only on how a test is built but also on who gets access to real test preparation and what conditions surround the testing itself, which makes the problem part psychometric and part social.
Practitioners cannot solve this by walking away from cognitive testing. When practitioners drop these tests for less structured methods, like unstructured interviews, adverse impact doesn't go away. It tends to move the problem onto tools that predict job performance less well to begin with. Kato et al. (2025) point out something uncomfortable: the way the field conceptualizes and measures cognitive ability at work has barely changed in a hundred years, even while neighboring fields in psychology and measurement have moved forward. Practitioners have to use century-old tools to meet legal and diversity standards that are entirely modern. Making sense of that mismatch, and what can actually be done about it, starts with understanding both halves of the dilemma in detail: what these tests measure and why they predict performance, and what adverse impact means and where it comes from.
What cognitive ability tests measure
Cognitive ability tests measure general mental ability, often shortened to GMA: the capacity to learn, reason through problems, and apply information in new situations. Decades of research back this construct as a predictor of performance across a wide range of jobs, from entry-level roles to highly complex technical and managerial positions.
Harver (2019) breaks GMA down into specific facets: reasoning, problem-solving, planning, abstract thinking, comprehension of complex ideas, and the ability to learn from experience. None of these facets are job-specific knowledge. They're general capacities that transfer across tasks and settings. Tests built to measure them usually take the form of verbal reasoning, numerical reasoning, logical or abstract reasoning, and spatial reasoning, and recruiters frequently combine several types to match the particular demands of a role.
GMA predicts job performance because of how people handle new information. When GMA is higher, employees tend to learn faster during onboarding, work through unfamiliar problems without a script, and carry knowledge from one task to another. Those capacities produce measurable performance differences, and the gap widens in roles with more complexity.
Recent research adds nuance to the simple GMA story. Kato et al. (2025) highlight work by Kato & Scherbaum (2023) on what they call ability tilt: a person's relative strength in one specific cognitive ability compared to another. Ability tilt predicts job performance beyond what overall GMA alone can explain, and its effect depends on whether the tilt lines up with what the job actually demands. A strong verbal tilt matters more in a writing-heavy role than it would in a role built around spatial reasoning. Kato et al. (2025) also report that two people with identical general intelligence scores can show very different patterns of specific abilities. One study identified five distinct cognitive profiles among people who scored similarly on overall GMA, which suggests that a single aggregate score can hide information that actually matters for whether someone fits a particular role.
Validity isn't flat across job types. Research cited in Campion & Campion (2025) confirms a complexity gradient: GMA predicts performance more strongly in cognitively demanding roles than in routine ones, where the ceiling on how much reasoning ability can help is lower to begin with.
The strength of the validity evidence
The validity evidence behind cognitive testing is substantial, but the size of the effect is genuinely disputed among researchers, and that dispute changes how cognitive tests should be weighed against other predictors. Nobody serious argues that GMA fails to predict job performance. The argument is over how much weight that predictive power deserves relative to other tools.
Schmidt & Hunter (1998), cited in Campion & Campion (2025), produced the meta-analytic finding that established GMA as one of the strongest predictors of job performance available to employers. That finding has held up remarkably well over time. A 2025 meta-analysis covering studies from 1949 to 2024 in Sweden reaffirmed GMA as a strong predictor of job performance, even after correcting for range restriction and measurement error in job performance ratings, confirming the pattern holds well outside North American samples where most of the earlier research originated.
Sackett et al. re-analyzed the data in 2022 and complicated the picture. Sackett et al. used different statistical corrections than Schmidt & Hunter did, and got an operational validity estimate well below the earlier figure, so you have to ask whether GMA's predictive power had been overstated for decades. Bobko et al. (2025) pushed back hard, arguing that the Sackett correction methodology is fundamentally flawed. That argument remains unresolved in the peer-reviewed literature as of this writing. For practitioners, the consequence is concrete: the case for treating cognitive tests as the single dominant predictor of job performance is weaker than it looked a decade ago, even though nobody disputes that cognitive tests have real predictive validity. The honest reading of the evidence falls between "cognitive tests work" and "cognitive tests are overrated. Validity is established. Its exact size, and its edge over other tools, is not settled science.
What adverse impact is
Adverse impact is defined by outcomes, not by what an employer intended. If a hiring test produces disproportionate results across protected groups, it can trigger legal liability even if it looks neutral and no one involved meant to discriminate.
eSkill (2025) defines adverse impact, also called disparate impact, as a condition where a hiring practice that appears neutral disproportionately harms members of a protected group, covering race, sex, age, and national origin, regardless of the employer's intent. Title VII of the Civil Rights Act has long been read to prohibit both intentional discrimination (disparate treatment) and practices that are neutral on paper but discriminatory in effect. That second category doesn't need proof of intent, and that is what makes adverse impact claims distinct from ordinary discrimination claims. As of June 2026, the DOJ's Office of Legal Counsel declared the EEOC's disparate-impact liability framework under Title VII unconstitutional, and this narrows federal enforcement of this doctrine going forward.
The operational standard most practitioners actually work against is the EEOC's four-fifths rule. Equalture (2022) and eSkill (2025) both describe it the same way: if a test's selection rate for a protected group falls below four-fifths of the selection rate for the group with the highest rate, the EEOC treats that gap as evidence of adverse impact and opens the test to review. The rule triggers scrutiny. It doesn't by itself establish that a practice is illegal.
Employers facing that scrutiny have a defense available, but it carries a heavy burden. eSkill (2025) explains that a cognitive test producing adverse impact does not automatically violate Title VII if the employer can show the test is valid, job-related, and that no equally effective but less discriminatory alternative exists. All three elements fall on the employer to prove, and the Dial Corp. case shows what happens when that proof fails. Dial's physical ability test disproportionately excluded women applicants, and when the EEOC challenged it, the court agreed that Dial had failed to prove a compelling need for the test and failed to show that non-discriminatory alternatives couldn't produce the same results. The court rejected the test's validity defense and found intentional discrimination.
The regulatory climate shifted again in April 2025. Executive Order 14281 directed federal agencies, including the EEOC, to deprioritize enforcement of disparate impact claims in favor of a merit-based approach to hiring oversight, eSkill reports. That order reduces enforcement pressure from federal regulators in the near term, but it changes nothing about underlying Title VII liability. Private plaintiffs retain the right to bring disparate impact claims on their own, independent of what the EEOC chooses to prioritize, and an executive order from one administration carries no guarantee of surviving the next. Treat this as a shift in enforcement climate, not a legal safe harbor. An employer that stops monitoring adverse impact because federal enforcement has eased is taking on risk it can't see coming, since private litigation doesn't depend on the EEOC's priorities.
Where group-score differences come from
The legal standard explains what triggers scrutiny. What produces the underlying score gaps in the first place is a separate question, and practitioners need an answer to it before they can judge whether, or how, those gaps can be narrowed. The group-score differences behind adverse impact claims are well documented and substantial in size, and some of them trace back to factors outside the cognitive construct being measured, so part of the gap may be addressable without giving up validity.
Campion & Campion (2025) confirm that cognitive tests produce larger subgroup differences than non-cognitive assessments like personality tests, structured evaluations of work experience, interest inventories, and reference checks. Among all the tools available to a hiring process, cognitive tests carry the largest documented group gaps. A meta-analysis of GMA testing in the United Kingdom, drawing on more than two million observations, adds weight to this picture from outside the American context. The scale of that dataset confirms that group differences on GMA tests aren't a byproduct of the particular social and educational history of the United States. The pattern is visible in a different country with a different educational system and a different demographic history.
The gap partly traces to unequal access to effective test preparation. Campion & Campion (2025) identify unequal preparation as one specific, documented mechanism: racioethnic minorities are more likely than non-minorities to use some forms of test preparation, but less likely to use the preparation tactics that are actually most effective. That finding matters because part of the score gap reflects differences in access to effective test-readiness resources, not just differences in underlying ability. Cultural and linguistic factors built into certain test formats have also been identified as sources of disadvantage, which raises the possibility that some test designs inflate apparent group differences beyond what the cognitive construct itself would produce on its own.
The research literature has not settled the deeper question of how much of the gap reflects test bias versus a genuine construct difference, and that debate isn't likely to close soon. What matters for a practitioner making decisions today is narrower and more useful: modern test designs applied to public safety occupations show much smaller score differences than traditional cognitive tests, according to Kato et al. (2025). That finding points to test modernization as a lever practitioners can actually pull, regardless of how the broader scientific debate over causation eventually resolves.
A newer population raises questions the literature hasn't caught up with yet. A 2025 study found almost no empirical evidence on whether the employment tests companies use regularly would reduce adverse impact for autistic applicants. The same study found that GMA tests, personality inventories, and situational judgment tests all produced meaningful subgroup differences unfavorable to autistic respondents. It remains unclear whether tools built around normative comparison or social judgment unintentionally disadvantage autistic individuals taking them. That's a gap in the research record practitioners should track as it develops, since the tools in common use today were not built or validated with this population specifically in mind.
What the FDNY case shows about real-world consequences
A large municipal fire department offers one of the clearest illustrations of how these dynamics play out when a cognitive or physical ability test, legal defense, and workforce diversity goals collide inside a single, high-profile hiring process. Entry exams for firefighter positions have a long history of producing exactly the kind of score gaps described above, and this department's hiring practices became the subject of sustained legal and public scrutiny over how those gaps affected minority applicants relative to white applicants.
The pattern in that case tracks the structural dilemma laid out from the start of this piece. A test designed to predict who could perform the job also produced the kind of group-score disparity that the four-fifths rule is built to catch, and the department carried the burden of defending the test's validity and job-relatedness under Title VII standards. That burden is what Dial Corp. failed to meet in the case described earlier. failed to meet in the case described earlier, and it's the same burden every employer using a high-stakes cognitive or physical test has to be ready to carry.
What reform can achieve in cases like this isn't the elimination of all group-score differences. Modern test design, better preparation access, and more careful attention to which specific abilities a role actually requires can narrow the gap that drives legal exposure without discarding the predictive power that makes cognitive testing valuable. Kato et al. (2025) already point to one concrete version of this: modern test formats applied to public safety roles produce meaningfully smaller subgroup differences than the traditional tests they replace. That's the direction the field is moving, and it's the most realistic answer available to the dilemma this piece opened with. Validity and fairness aren't fully reconcilable through a single policy change, but the gap between them is not fixed, either. It responds to deliberate, specific work on test design and preparation access, and that work is where practitioners actually have room to act.
Sources
- Using Practice Employment Tests in Recruitment and Selection to Equalize Preparation Opportunities - Campion - 2025 - Human Resource Management - Wiley Online Library
- Cognitive Ability Testing in the Workplace: Modern Approaches and Methods
- Full article: General mental ability testing and adverse impact in the United Kingdom: a meta-analysis with more than two million observations
- General Cognitive Ability and job performance in personnel selection in Sweden: A meta-analysis - Anders Sjöberg, Sofia Sjöberg, 2025
- United States V. City Of New York FDNY Employment Discrimination Case


