Skip to main content
AssessIQ Sign in

Glossary · Assessment Science

Reliability Coefficient

A reliability coefficient is a number between 0 and 1 that expresses how consistently a test measures the same thing — across repeated administrations, equivalent test forms, or multiple raters. A coefficient of 1.0 would mean perfect consistency with zero measurement error; 0 would mean pure noise. In high-stakes hiring, a coefficient above 0.80 is the widely cited minimum standard.

Why it matters in hiring and assessment.

Reliability is a precondition for validity. A test that gives different results each time a candidate takes it — not because the candidate changed, but because of random measurement error — cannot validly predict anything. Statistically, the maximum possible correlation between a test and a criterion (validity) is limited by the square root of the test's reliability coefficient. A test with reliability of 0.64 can achieve a maximum validity of 0.80 — and real-world values are almost always lower due to criterion unreliability too.

There are three main types of reliability, each measured differently:

  • Test-retest reliability: The same candidates take the test twice; the correlation of the two sets of scores is the coefficient. Sensitive to practice effects and the time interval chosen.
  • Internal consistency (e.g. Cronbach's alpha): Measures whether all items in a test correlate with one another — useful for ability and knowledge tests. Does not require two administrations. Can be artificially inflated by a very large number of items or by item redundancy.
  • Inter-rater reliability: The degree to which two or more scorers independently assign the same scores to the same responses — critical for structured interviews, coding exercises evaluated by humans, or essay-based questions. Cohen's kappa and intra-class correlation (ICC) are common metrics.

For automated, objectively scored assessments (MCQ, coding test-case scoring), internal consistency reliability is usually high and reproducible. For subjectively scored responses — structured interview ratings, written answers, open-ended coding rubrics — inter-rater reliability requires explicit anchor design and rater calibration to stay above the defensible threshold.

Example.

An aptitude test vendor reports Cronbach's alpha = 0.87 for their numerical reasoning battery. This means approximately 87 % of score variance reflects the underlying construct (numerical reasoning), and 13 % is measurement error. A candidate who scored 72 on this test would, in theory, score between 65 and 79 on repeated equivalent-form testing — the standard error of measurement quantifies this band. At 0.87, the test meets the common 0.80 threshold, and a validity ceiling of √0.87 ≈ 0.93 gives ample room for meaningful criterion prediction.

  • Criterion Validity

    The observed criterion validity coefficient is limited by the reliability of both the test and the criterion measure — a fundamental psychometric constraint.

  • Construct Validity

    Reliability is necessary but not sufficient for construct validity — a reliable test can consistently measure the wrong thing.

  • Item Response Theory

    IRT provides item-level reliability information (information functions) rather than a single global coefficient, enabling adaptive tests to target measurement precision where it matters most.

  • Cut Score

    The reliability coefficient determines the width of the standard error of measurement band around a cut score — low reliability means more candidates near the cut could legitimately be above or below it.