Skip to main content
AssessIQ Sign in

Glossary · Assessment Science

Item Response Theory

Item Response Theory (IRT) is a family of psychometric models that describe the probability of a correct response as a mathematical function of the test-taker's underlying ability and the characteristics of the specific item — principally its difficulty and how well it discriminates between ability levels. Unlike classical test theory, IRT ability estimates are independent of which specific items were administered, making them comparable across test forms and candidates.

Why it matters in hiring and assessment.

Classical test theory (CTT) — the older framework — scores candidates by summing correct answers and calibrates items based on how difficult they were for the specific sample that took them. This means a "hard" item is only hard relative to the particular group tested, and two candidates who sit different versions of a test cannot be fairly compared without statistical equating. IRT solves both problems by modelling difficulty and ability on the same latent scale, independent of the sample.

In applied hiring assessment, IRT matters for three reasons:

  • Adaptive testing: IRT is the mathematical engine behind computer-adaptive testing (CAT). Because the model can estimate ability from any subset of calibrated items, it can select the next question based on real-time ability estimates — making the test shorter and more precise simultaneously.
  • Item banking and form equivalence: Large-scale assessors maintain banks of IRT-calibrated items and can assemble new test forms that are statistically equivalent, preventing score inflation as items become exposed.
  • Differential Item Functioning (DIF) analysis: IRT enables detection of items that are unexpectedly harder or easier for specific demographic subgroups after controlling for ability — a key tool in adverse impact investigation and test fairness review.

The most common IRT models are the 1PL (Rasch), 2PL, and 3PL, differing in whether they model item discrimination and a lower-asymptote guessing parameter in addition to difficulty. The Rasch model is particularly favoured in high-stakes certification because its strict requirements — when met — guarantee that ability estimates are person-free and item-free.

Example.

A platform maintains a bank of 500 IRT-calibrated SQL questions. Candidate A is administered 30 questions drawn randomly; candidate B takes a different 30. Because both sets are calibrated on the same latent scale, their ability estimates (expressed as theta, θ) can be directly compared even though they answered different items. A candidate with θ = 1.5 is estimated to have a 75 % probability of correctly answering an item at the same difficulty level — the item characteristic curve makes this explicit and auditable.

  • Computer-Adaptive Testing

    CAT is the applied implementation of IRT — adaptive item selection requires a calibrated IRT model to function.

  • Reliability Coefficient

    IRT expresses reliability as a test information function rather than a single coefficient, showing where on the ability scale measurement is most precise.

  • Construct Validity

    IRT's DIF analysis is a tool for detecting items that violate construct validity by behaving differently across demographic groups after controlling for ability.