Skip to main content
AssessIQ Sign in

Glossary · Assessment Science

Criterion Validity

Criterion validity is the statistical relationship between scores on a selection test and a real-world performance criterion — typically supervisor ratings, output metrics, or training results. It is expressed as a correlation coefficient (r). In personnel selection, values above 0.30 are generally considered practically significant, meaning the test adds meaningful predictive power over chance.

Why it matters in hiring and assessment.

Criterion validity is the most direct evidence that a test is worth using in hiring. It answers the question: does a higher score actually predict better performance on the job? Without this evidence, a test may feel rigorous and look professional while systematically selecting for traits that do not drive results. Structured skills tests, coding assessments, and aptitude batteries are defensible precisely because their criterion-related validity can be (and has been, in the research literature) demonstrated empirically.

There are two subtypes. Predictive validity is established by testing candidates before hire, then correlating their scores with later performance data — the gold standard because it matches the actual selection context. Concurrent validity tests current employees and correlates their scores with existing performance ratings — faster to obtain but potentially biased because poor performers may already have been separated and top performers may score higher due to job experience rather than pre-existing ability.

When an organisation is challenged on a hiring decision — or faces an adverse impact claim — criterion validity evidence is one of the primary defences. A vendor who cannot produce validity data (ideally a peer-reviewed study or a client-commissioned local validation) is selling on brand, not science.

Example.

A company administers a Python coding test to 120 developer applicants and hires 40 of them. Twelve months later, it correlates each hire's test score with their 12-month performance rating. The resulting correlation is r = 0.42, which is above the 0.30 threshold and statistically significant at p < 0.01. This is criterion validity evidence. The company can now state that higher test scores predicted stronger on-the-job performance in this role — making the test a defensible component of the selection process.

  • Construct Validity

    Evidence that the test measures the psychological attribute it is designed to measure — a prerequisite for interpreting criterion relationships meaningfully.

  • Reliability Coefficient

    A test cannot be more criterion-valid than it is reliable — measurement error in the test attenuates the observed validity coefficient.

  • Adverse Impact

    Criterion validity is the primary evidence of business necessity used to justify a selection tool that produces disparate pass rates.

  • Cut Score

    Setting a defensible cut score typically requires criterion validity data to determine where on the score scale job-relevant ability becomes sufficient.