Glossary · Assessment Science
Reliability Coefficient
A reliability coefficient is a number between 0 and 1 that expresses how consistently a test measures the same thing — across repeated administrations, equivalent test forms, or multiple raters. A coefficient of 1.0 would mean perfect consistency with zero measurement error; 0 would mean pure noise. In high-stakes hiring, a coefficient above 0.80 is the widely cited minimum standard.
Why it matters in hiring and assessment.
Reliability is a precondition for validity. A test that gives different results each time a candidate takes it — not because the candidate changed, but because of random measurement error — cannot validly predict anything. Statistically, the maximum possible correlation between a test and a criterion (validity) is limited by the square root of the test's reliability coefficient. A test with reliability of 0.64 can achieve a maximum validity of 0.80 — and real-world values are almost always lower due to criterion unreliability too.
There are three main types of reliability, each measured differently:
- Test-retest reliability: The same candidates take the test twice; the correlation of the two sets of scores is the coefficient. Sensitive to practice effects and the time interval chosen.
- Internal consistency (e.g. Cronbach's alpha): Measures whether all items in a test correlate with one another — useful for ability and knowledge tests. Does not require two administrations. Can be artificially inflated by a very large number of items or by item redundancy.
- Inter-rater reliability: The degree to which two or more scorers independently assign the same scores to the same responses — critical for structured interviews, coding exercises evaluated by humans, or essay-based questions. Cohen's kappa and intra-class correlation (ICC) are common metrics.
For automated, objectively scored assessments (MCQ, coding test-case scoring), internal consistency reliability is usually high and reproducible. For subjectively scored responses — structured interview ratings, written answers, open-ended coding rubrics — inter-rater reliability requires explicit anchor design and rater calibration to stay above the defensible threshold.
Example.
An aptitude test vendor reports Cronbach's alpha = 0.87 for their numerical reasoning battery. This means approximately 87 % of score variance reflects the underlying construct (numerical reasoning), and 13 % is measurement error. A candidate who scored 72 on this test would, in theory, score between 65 and 79 on repeated equivalent-form testing — the standard error of measurement quantifies this band. At 0.87, the test meets the common 0.80 threshold, and a validity ceiling of √0.87 ≈ 0.93 gives ample room for meaningful criterion prediction.
Related terms.
- Criterion Validity
The observed criterion validity coefficient is limited by the reliability of both the test and the criterion measure — a fundamental psychometric constraint.
- Construct Validity
Reliability is necessary but not sufficient for construct validity — a reliable test can consistently measure the wrong thing.
- Item Response Theory
IRT provides item-level reliability information (information functions) rather than a single global coefficient, enabling adaptive tests to target measurement precision where it matters most.
- Cut Score
The reliability coefficient determines the width of the standard error of measurement band around a cut score — low reliability means more candidates near the cut could legitimately be above or below it.