Skip to main content
AssessIQ Sign in

Glossary · Assessment Science

Computer-Adaptive Testing

Computer-adaptive testing (CAT) is an assessment format in which the difficulty of each successive item is chosen in real time based on the test-taker's estimated ability, updated after every response. Rather than presenting every candidate with the same fixed question set, the algorithm targets items near the estimated ability level — concentrating measurement where it is most informative and reaching a precise score estimate with significantly fewer items.

Why it matters in hiring and assessment.

Fixed-form tests waste questions on both ends of the ability distribution. A very strong candidate answers the easy items correctly with certainty, gaining almost no measurement information. A weaker candidate answers the hardest items incorrectly with near certainty, also gaining little information. CAT eliminates most of this waste by bypassing items that are too easy or too hard for a given candidate and concentrating on items at the edge of their ability — the zone where response uncertainty is highest and information is greatest.

The practical benefits for candidate screening are concrete:

  • Shorter tests, same precision: CAT typically reaches equivalent measurement precision with 40–60 % fewer items than a fixed-form test. A 60-item aptitude battery can be reduced to 25–35 adaptive items without sacrificing the accuracy of the ability estimate.
  • Reduced floor and ceiling effects: Fixed forms often fail to discriminate among candidates at the top or bottom of the score range. Adaptive selection avoids this by continuously targeting the ability frontier.
  • Better candidate experience: Candidates are not grinding through items far below their level (demoralising) or consistently failing items far above it (discouraging). The adaptive difficulty feels more like a conversation than a gauntlet.
  • Security: Because no two candidates see exactly the same item sequence, memorising and sharing answers provides limited advantage compared with a fixed form where the full item set can be reconstructed and leaked.

CAT requires a calibrated item bank — typically at least 200–500 items per domain calibrated via Item Response Theory — before it can function. This makes it an infrastructure investment not suited to small-scale or bespoke assessment programmes.

Example.

A platform launches a 40-item CAT aptitude test. Candidate A answers the first medium-difficulty item correctly; the algorithm updates the ability estimate upward and serves a harder item. After five correct answers in a row, the estimated ability stabilises in the 85th-percentile range and easier items are no longer administered. The test terminates after 22 items once the standard error of measurement falls below a preset threshold. Candidate B, answering inconsistently, takes the full 40 items because the algorithm cannot reduce uncertainty quickly. Both receive ability estimates on the same scale, comparable against the same norm group.

  • Item Response Theory

    The psychometric model that powers adaptive item selection — CAT cannot function without a calibrated IRT item bank.

  • Reliability Coefficient

    CAT targets a reliability threshold (expressed as a standard error bound) as its stopping criterion, rather than a fixed item count.

  • Percentile Rank

    CAT scores are typically reported as percentile ranks against a norm group, translating the IRT theta score into a hiring-practical comparison point.