Are Pre-Employment Assessments Legally Risky? What the Research and Case Law Actually Say

Why validated hiring assessments can offer a more consistent, measurable, and legally defensible approach to employee selection.

‍

Introduction: The Hiring Risk You May Not Be Measuring

‍

Almost every week, I talk with an HR leader who sees the value in pre-employment assessments but hesitates to introduce them. The concern often comes from legal. What if the assessment creates adverse/disparate impact? What if a personality test raises ADA concerns? What if candidates from different racial groups receive different scores?

‍

These are legitimate questions. They deserve serious answers. But there is another question I think employers should be asking: What is the legal risk of continuing to make hiring decisions without a standardized, consistent, and validated process?

‍

Choosing not to use an assessment does not eliminate hiring bias. It usually means relying more heavily on resumes, interviews, and manager judgment. Those methods can produce discriminatory outcomes too, and they are harder to defend when the criteria were never written down.

‍

The answer is not to replace hiring managers with talent assessments and skills tests. It is to create a hiring process in which every decision is supported by relevant evidence. That is where well-designed, validated hiring assessments can make a meaningful difference.

‍

TL;DR: Key Takeaways

‍

·       Pre-employment assessments can support more consistent, evidence-based hiring when they are job relevant, appropriately validated, and used with fairness monitoring.

·       Avoiding formal assessments does not eliminate hiring bias. Interviews and manager judgment also need clear criteria and documentation.

·       Cognitive tests can predict performance while producing demographic score differences. Employers should evaluate job relevance, cutoffs, and alternatives.

·       Personality assessments of ordinary work behavior differ from clinical tests designed to identify mental disorders. Purpose and design matter.

·       The 80% rule is a screening tool, not a guarantee of fairness or a finding of discrimination. Examine sample size, statistical evidence, and each selection stage.

‍

1. The Hiring Bias You’re Not Measuring: Interviews and Human Judgment

‍

Most organizations would never allow a hiring manager to administer an unvalidated personality test they created themselves. Yet many allow that same manager to make a hiring decision after an informal conversation with few predetermined questions, inconsistent evaluation criteria, and little documentation. Why do we scrutinize one process so carefully while assuming the other is safe?

‍

Each interview is a selection assessment, too

‍

The Supreme Court settled this in Watson v. Fort Worth Bank & Trust (1988). Clara Watson, a Black employee, was passed over for promotion four times by supervisors relying on their own judgment. The Court held that disparate-impact analysis applies to subjective practices exactly as it applies to standardized tests. Avoiding formal tests does not put an employer outside discrimination law.

‍

What does research say about structured interviews?

Sackett, Zhang, Berry, and Lievens (2022) reassessed roughly a century of selection research and corrected statistical problems that had inflated earlier estimates. Their work gives employers two numbers for each method: how well it predicts job performance, and how large an average gap it produces between Black and White applicants.

‍

Selection method Validity for job performance Black-White difference (d)
Interviews
Structured interview .42 .23
Unstructured interview .19 .32
Ability and judgment
Cognitive ability test .31 .79
Situational judgment test .26 .39
Personality, work-context
Conscientiousness .25 -.07
Emotional stability .23 .09

‍

Source: Sackett, Zhang, Berry, and Lievens (2022). Validity is the correlation with job performance. A larger d means a wider average gap; a negative value means the second group scored slightly higher. Sample types vary behind the d estimates, so the ordering is more dependable than exact values.

‍

Three things stand out. The structured interview predicts better than anything else listed while producing one of the smaller gaps, so the most rigorous method here is not the riskiest one. The cognitive ability test produces the largest gap by a wide margin but still predicts well. The two personality measures show the smallest subgroup differences in the table, with conscientiousness slightly favoring Black applicants, while predicting about as well as a cognitive test. Work-context matters for those rows: items asking about behavior at work rather than in general predict better. Sackett and colleagues caution that these figures are not a mandate for preferring one method regardless of circumstances (Sackett et al., 2023). Job relatedness should be the determining factor.

‍

For more on separating candidate presentation from performance evidence, see our guide to pre-employment assessments in the age of AI.

 

2. Are Pre-Employment Assessments Legal? What the Case Law Actually Says

‍

The history of employment testing law is often told as a warning against assessments. Read closely, it says something more useful. Employers need to know what their selection procedures measure, why those characteristics matter for the work, and how the results are used. The same standard applies to interviews, experience requirements, and promotion decisions, not only to tests.

‍

Decision Practical relevance
Griggs v. Duke Power Co. (1971) Facially neutral requirements can violate Title VII when they produce discriminatory effects and lack a sufficient work-related justification.
Albemarle Paper Co. v. Moody (1975) Validation evidence must meaningfully relate to the work; merely presenting a validation study is not enough.
Connecticut v. Teal (1982) Favorable overall hiring results do not automatically excuse a discriminatory intermediate selection step.
Watson v. Fort Worth Bank & Trust (1988) Subjective selection practices can be challenged under disparate-impact analysis.
Ricci v. DeStefano (2009) Discarding selection results because of racial disparities can itself create legal problems without a sufficient evidentiary basis.

‍

What changed in federal enforcement in 2025 and 2026?

‍

Federal policy shifted significantly. Executive Order 14281, signed April 23, 2025, directed federal agencies to deprioritize disparate-impact enforcement. On June 9, 2026, the Department of Justice’s Office of Legal Counsel issued an opinion challenging the EEOC’s historical interpretation of disparate-impact liability and associated validation requirements. Rescission of the Uniform Guidelines on Employee Selection Procedures also appeared on the EEOC’s regulatory agenda.

‍

An executive order or agency legal opinion does not repeal Title VII's statutory disparate-impact provision, 42 U.S.C. § 2000e-2(k). Employers and counsel should distinguish enforcement priorities, agency interpretations, statutory obligations, and binding court decisions.

‍

3. What Makes Pre-Employment Assessments Validated and Defensible?

‍

An assessment is not defensible because a vendor calls it scientific. Evidence has to support the conclusions drawn from scores and the decisions made with them. Our approach at ForPsyte rests on three activities: understanding the job, examining fairness, and connecting results to performance.

‍

Start with job analysis: What does success actually require?

‍

Before choosing an assessment, identify the work. What separates an effective restaurant general manager from an average one? What judgment does a bank teller need when a compliance concern comes up? Job analysis identifies the relevant tasks and worker characteristics, which gives employers a rationale for what an assessment measures.

‍

It also stops employers from buying a test because its label sounds impressive. A broad cognitive test does not belong in a process simply because intelligence correlates with performance in general. Job analysis is what establishes job relatedness.

‍

Like a baseball team checking which statistics predict wins, employers need to test which hiring signals predict success. Our Moneyball for hiring article explains this approach to validation.

‍

Evaluate fairness: What happens across groups?

‍

Examine whether a procedure produces meaningfully different outcomes across demographic groups, including the assessment, the scoring, the cutoffs, and how results feed later decisions. As the table in section 1 shows, validity and fairness are separate questions, and both need answers.

‍

Connect pre-employment assessment scores to actual performance

‍

Criterion-related validation examines whether scores relate to outcomes such as supervisor ratings, sales, service, productivity, or retention. A restaurant group can assess its current managers, gather performance ratings, and see which characteristics actually track with performance. Those results often reveal that an appealing trait adds little.

‍

A local study is particularly useful because it tests the question in the employer's own setting, though professional standards recognize other sources of validity evidence. What matters is whether the evidence supports the specific proposed use (SIOP, 2018).

‍

For a practical example, see our restaurant leadership assessment case study, which connects manager assessment results to supervisor performance ratings.

‍

4. Cognitive Ability Tests and Racial Differences: The Hardest Assessment Question

‍

Cognitive ability testing worries general counsel because research documents average score differences between racial and ethnic groups, and depending on the applicant pool and cutoff those differences can produce adverse impact. That does not make the testing unlawful. It means employers need evidence for what they measure and how they use it.

‍

How large are the reported differences?

‍

The .79 difference in the table above is the largest of any method listed. If a cutoff passes 50% of the higher-scoring group, roughly 21% of the other group would pass, an impact ratio near 42% and far below the 80% threshold. Actual outcomes depend on the test, the applicant pool, and the cutoff.

‍

Does cognitive ability still predict performance?

‍

Yes, but newer research has challenged how much predictive value it adds.

‍

Schmidt and Hunter (1998) put its validity at .51. The .31 above reflects Sackett et al.'s (2022) revision after reconsidering earlier statistical corrections, a revision researchers still debate (Bobko et al., 2025). Berry et al. (2024) then found that removing cognitive tests from many assessment combinations substantially reduced adverse impact with little or no loss in predicted performance.

‍

These tests predict best in complex roles demanding more learning, reasoning, and problem solving (Hunter & Hunter, 1984), which makes them a better fit at management and leadership levels than across the board.

‍

What do cognitive-testing lawsuits teach us?

‍

Three cases show why design and implementation decide the outcome. Lewis v. City of Chicago (2010) involved repeated hiring from a firefighter list using a cutoff higher than the announced passing score. United States v. City of New York (2009) challenged written firefighter exams used to screen and rank applicants. EEOC v. Ford Motor Co. (2005) settled a claim over an apprenticeship test. Job relevance, cutoffs, ranking practices, and available alternatives all shaped the risk.

‍

How should employers use cognitive testing?

‍

·       Align the use of testing to role complexity.

·       Select relevant cognitive measures.

·       Establish and justify cutoffs using job-related evidence.

·       Consider whether combining cognitive tests with structured interviews, personality, and situational judgment assessments provides comparable prediction with less adverse impact.

·       Monitor selection rates and predictive validity as jobs and applicant populations change.

‍

The goal is not to avoid cognitive ability when it matters. It is to measure the right attributes, in the right combination, for the right reasons.

‍

5. Personality Assessments and the ADA: Are Employers Testing Mental Health?

‍

The other objection I hear concerns the Americans with Disabilities Act. Can an employer assess emotional stability? Do questions about mood or stress become an unlawful pre-offer medical examination? The distinction is between measuring normal-range work behavior and using an instrument designed to identify mental disorders.

‍

Are personality tests medical examinations?

‍

Not automatically. The ADA prohibits pre-offer medical examinations and disability-related inquiries. EEOC guidance separates nonmedical measures of ordinary traits, preferences, and work habits from tests built to reveal mental impairment. Whether an assessment counts as medical depends on its purpose, design, content, administration, and interpretation, not the vendor's label.

‍

Why Karraker v. Rent-A-Center matters

‍

In Karraker v. Rent-A-Center (2005), employees challenged the use of the Minnesota Multiphasic Personality Inventory in a promotion process. The Seventh Circuit held that this clinical instrument was a medical examination even though the employer scored it for vocational purposes. The lesson is that employers have to know what an instrument was designed to reveal.

‍

Can emotional stability be assessed?

‍

Emotional stability is a normal-range trait covering tendencies such as staying composed and managing frustration, and the table in section 1 puts work-context measures at .23 validity with a subgroup difference of .09.

‍

Questions about psychiatric symptoms raise different issues than questions anchored in ordinary workplace behavior, though wording alone is not a safe harbor. Ask providers how they define the construct and what evidence supports the intended job use.

‍

What about accommodations?

‍

Have a process for reasonable accommodation requests covering an assessment's format, accessibility, or administration. Standardization is not a reason to refuse accommodations. The objective is to measure job-related characteristics without adding barriers unrelated to the work, and individual requests need individual review.

‍

6. What Is Adverse Impact, and How Should Employers Measure It?

‍

Adverse impact happens when a hiring practice selects one demographic group at a substantially lower rate than another, even though everyone went through the same process. The usual check is the four-fifths rule, also called the 80% rule. Take the share of each group that gets through a step, find the highest share, and divide the others by it. Anything below 80% marks that step for a closer look. Under the historical Uniform Guidelines this is a screening tool, not a finding of discrimination.

‍

Where the 80% line actually falls

‍

Picture two separate hiring steps, each a test or screening interview that applicants either pass or do not, and each given to 400 people split evenly into two groups of 200. The results land one applicant apart, and that one applicant decides whether the rule flags the step.

‍

Comparison of two hiring examples: an 80% gender selection-rate ratio and a 79% age selection-rate ratio, illustrating the limits of a strict threshold.

Example 1: Gender differences that meet the 80% threshold

‍

Men passed the first step at 50.0%, women at 40.0%, an impact ratio of exactly 80%. That meets the threshold, so the rule does not flag it and an employer scanning a spreadsheet moves on. That is the trap. A two-proportion test on the same numbers returns a z of 2.01, roughly where the Supreme Court in Castaneda v. Partida and Hazelwood School District v. United States began treating statistical gaps as meaningful.

‍

Example 2: Age differences that warrant further investigation

‍

Applicants under 40 passed at 50.0%, applicants 40 or older at 39.5%, an impact ratio of exactly 79%. One applicant separates this from the example above. Same group sizes, same reference rate, one fewer person through the door, and the outcome flips from no flag to flagged. Its z value is 2.11, almost identical to the first. The rule changed its answer. The evidence did not.

‍

Important distinction

‍

The historical Uniform Guidelines apply the four-fifths rule to race, sex, and ethnic groups, not to age. The age example above borrows the same arithmetic as an illustration, not as a formal rule under the ADEA. That statute protects applicants 40 and older under a different framework, where an employer defends by showing a reasonable factor other than age (Smith v. City of Jackson, 2005).

‍

What should employers do when they identify a disparity?

‍

A result on either side of the line is a prompt to look, not a conclusion. Start with sample size, because Roth, Bobko, and Switzer (2006) found the rule produces frequent false alarms in small samples, which is why courts pair it with a significance test. Check each stage rather than only who gets hired, since Connecticut v. Teal established that a step which screens people out has to stand on its own. Then investigate the cause and find a job-related alternative with less impact.

‍

One boundary matters more than the rest. Monitoring is not permission to adjust someone's score or move a cutoff because of race, sex, or age. Ricci v. DeStefano made that an expensive lesson for New Haven.

‍

The goal is not simply to pass the 80% rule. It is to build a hiring process that predicts job performance, gives people a fair shot, and can be supported by evidence.

 

7. Six Questions to Ask Your Assessment Provider

‍

1.          How do you establish job relatedness? How were the competencies identified, why do they matter for this role, and was the assessment linked to the job through job analysis?

2.          What evidence supports predictive validity? Are scores associated with meaningful performance outcomes, and is that evidence relevant to our proposed use?

3.          What do we know about demographic differences? Has adverse impact been examined, and what limitations apply to the data?

4.          How are scores and cutoffs established? Is the approach evidence-based rather than an arbitrary passing score?

5.          How are ADA considerations addressed? Is the instrument nonmedical, and what is your approach to reasonable accommodation?

6.          What happens after implementation? Can the provider monitor fairness, evaluate local outcomes, and revisit the assessment as the job changes?

‍

A provider should answer these with documentation rather than marketing claims. If it cannot explain what the tool measures or why the scores support a hiring decision, settle that before implementation.

‍

Use the assessment provider guide and checklist in our customer resources to organize those questions and compare the evidence providers supply.

‍

Frequently Asked Questions

‍

Are pre-employment assessments legal?

‍

Yes. Employers can use pre-employment assessments, but their design and use must comply with applicable law. Job relevance, validity evidence, appropriate administration, accommodations, and monitoring help employers support their selection decisions.

‍

Do pre-employment assessments eliminate hiring bias?

‍

No assessment guarantees a bias-free hiring process. Standardized measures can make decisions more consistent and easier to evaluate, but employers still need to examine outcomes and how scores influence each hiring stage.

‍

Does meeting the 80% rule prove a hiring process is fair?

‍

No. The four-fifths rule is a screening tool. Meeting its threshold does not rule out a meaningful disparity, and falling below it does not by itself establish unlawful discrimination.

‍

Can personality assessments be used before a job offer?

‍

Nonmedical assessments of ordinary work-related traits may be used before a job offer. Tests designed to reveal mental impairment raise different ADA concerns. Review the instrument’s purpose, content, and interpretation, rather than relying on its label.

‍

8. A More Defensible Hiring Decision Starts With Better Evidence

‍

When an HR leader tells me legal is worried about assessments, I understand why. Cognitive testing can produce substantial demographic differences. Clinical instruments can raise real ADA problems. Unsupported scoring, arbitrary cutoffs, and sloppy administration all create exposure. Those risks are real, and so are the risks of doing nothing.

‍

A manager's judgment is not objective because it comes from experience, and an interview is not fair because everyone had a conversation. A familiar process is not defensible because it has been in place for years.

‍

A well-designed assessment gives you defined measures, consistent criteria, and evidence you can examine and improve. Supported by job analysis, validity evidence, fairness evaluation, accommodations, and monitoring, it helps employers decide better and manage risk. That is not a guarantee against challenges. It is a commitment to evidence over assumption.

‍

Want to know whether your current hiring assessments are supported by the right evidence? Talk with ForPsyte about your selection process, validity evidence, and fairness monitoring.

‍

Legal note: This article offers general information about U.S. employment selection practices, not legal advice. Consult qualified employment counsel about specific circumstances and applicable federal, state, and local requirements.