Assessment Methodology
How AssessIQ assessments are designed and scored.
A hiring assessment is only useful if it measures what it claims to measure, consistently, and without unfair disadvantage to any group of candidates. This page describes the principles and practices behind AssessIQ question packs, the band-scoring model, and our approach to fairness and integrity.
A note on formal validation: AssessIQ describes design principles and internal quality-control practices — not the results of formal psychometric validation studies. We do not hold I-O psychology certifications, nor have we published peer-reviewed validity or reliability data for our question packs. We reference established standards (SIOP Principles, EEOC Uniform Guidelines, APA/AERA/NCME Standards) to describe the principles our design process aims to follow, not to claim compliance with those standards. Where we cannot honestly make a claim, we say so.
Question design
How question packs are built.
Job-task analysis first
Questions start from what the job actually requires.
Before a question pack is written, the relevant competency domain is defined: what does a candidate in this role need to know or be able to do? Questions are scoped to observable, job-relevant skills — not academic trivia or recall of syntax that any IDE would supply. This practice reflects the job-relatedness requirement at the centre of the EEOC Uniform Guidelines on Employee Selection Procedures (1978).
Item-writing principles
Each question is written to a standard, not a feeling.
Items follow established item-writing guidelines: one defensible correct answer per MCQ; stems that are complete and unambiguous; distractors that are plausible but clearly distinguishable from the correct answer by someone with the target skill; and language that does not require specialised vocabulary unrelated to the competency being measured. Trick questions and double negatives are explicitly disallowed.
Internal review before deployment
No question is deployed after a single author’s judgment.
Every question pack goes through a review pass before it is made available: the reviewer checks item quality, language neutrality, and whether the anchor descriptors (the rubric language) accurately describe what a response at each band would look like. This is a human-in-the-loop gate, not an automated approval. We do not claim this constitutes a formal psychometric validation study — it is a quality-control step.
Question-pack versioning
Packs are versioned; scoring is tied to the version used.
When a question pack is updated, the prior version remains available for assessments already in progress. Every attempt record carries the pack version and question set used, so scoring and results are reproducible and auditable against the exact content the candidate saw.
Scoring model
The five-band scoring scale.
AssessIQ scores every response on a five-point ordinal band: 0, 25, 50, 75, 100. These are not percentages — they are discrete levels of demonstrated competence, each anchored to a description of what a response at that level looks like. A score of 50 means “functional, working-level skill demonstrated,” not “half-correct.” See cut-score for how organisations translate band scores into pass/fail thresholds.
Not demonstrated
The response does not demonstrate the target competency. The answer is absent, fundamentally incorrect, or shows a misconception that would cause harm in a working context.
Partial / foundational
Some relevant knowledge or skill is visible, but with significant gaps. The candidate is aware of the concept but cannot apply it reliably or completely.
Functional
The target skill is demonstrated at a working level. The response is mostly correct with minor gaps or imprecision that would not block a reasonably supervised practitioner.
Proficient
The skill is demonstrated with confidence and accuracy. The response goes beyond the minimum — considering edge cases, trade-offs, or best practices — in ways that distinguish a proficient practitioner from a functional one.
Expert
The response reflects mastery: complete, precise, and showing the depth of understanding that would allow the candidate to teach or extend the concept, not merely apply it.
Reliability & validity
What we can and cannot claim.
The psychometric standards most relevant to employment assessment are the SIOP Principles for the Validation and Use of Personnel Selection Procedures, the EEOC Uniform Guidelines on Employee Selection Procedures, and the APA/AERA/NCME Standards for Educational and Psychological Testing. We reference these to name the concepts our design aims to honour — not to assert formal compliance.
Construct validity (design intent)
A test has construct validity when its scores actually reflect the underlying skill it claims to measure, not something else (test-taking ability, guessing, cultural familiarity with question style). AssessIQ aims for construct validity by grounding every question in observable job-task behaviours and by writing anchor rubrics that describe skill directly — not effort, confidence, or writing ability (except where those are the competency). We cannot claim formally established construct validity in the psychometric sense (which requires administration to a reference population and statistical analysis) — we can only claim that our design process is consistent with the construct-validity principle as described in the APA/AERA/NCME Standards for Educational and Psychological Testing.
Criterion-related validity (the honest limit)
Criterion-related validity — whether assessment scores predict actual job performance — requires longitudinal data: correlating scores with supervisor ratings or business outcomes after hire. AssessIQ does not have that data for your organisation. No pre-built assessment platform can provide criterion validity evidence for your specific role and context without running that study in your hiring pipeline. What we provide is a structured, consistent signal. Whether it predicts performance in your context is something you can test by tracking outcomes — and we encourage you to do so. See the /glossary/criterion-validity entry for more on what this means in practice.
Reliability (consistency)
A reliable assessment produces consistent results: a candidate who takes the same assessment on two occasions, or two candidates of identical ability, should score similarly. AssessIQ supports reliability through standardised question text, standardised rubrics with explicit band anchors, and time-limited conditions that are enforced server-side. The AI-assisted scoring layer is calibrated to the same rubric a human reviewer would apply. We do not publish an internal reliability coefficient for our question packs because we have not run the reference-population studies required to calculate one honestly. See /glossary/reliability-coefficient for what that calculation requires.
Related glossary terms
Fairness
Reducing unfair disadvantage.
Language neutrality review
Question packs are reviewed for language that may disadvantage candidates based on regional dialect, cultural context, or educational background unrelated to the target skill. Items that test English fluency when the role does not require it, or that use idioms specific to one cultural context, are flagged and rewritten.
Job-relevance gate
The EEOC Uniform Guidelines on Employee Selection Procedures require that selection procedures be job-related. Every question in AssessIQ is scoped to a defined competency domain with an explicit job-relevance rationale. Questions that cannot be tied to a job task are excluded.
Band-score transparency
Every scored response carries a rationale: what the candidate demonstrated, and at which band. This makes the basis for a score inspectable — by administrators, by internal auditors, and if challenged, by the candidate. Opaque pass/fail scores with no explanation are a fairness and legal-exposure risk; AssessIQ is designed to avoid them.
Adverse-impact awareness
Adverse impact occurs when a selection procedure produces significantly different pass rates across demographic groups. AssessIQ does not collect candidate demographic data, so we cannot calculate adverse-impact ratios on your behalf. Organisations running large-volume hiring drives are encouraged to track pass rates by group independently and to review question content if group differences emerge. See /glossary/adverse-impact for the four-fifths rule and how to apply it.
Proctoring integrity
Conditions that make scores meaningful.
A valid score requires a controlled test environment. AssessIQ’s proctoring layer — available on Growth and Enterprise tiers — combines browser lockdown (candidates cannot switch tabs or open other applications) with webcam monitoring and server-side time enforcement. See our proctoring glossary entry and the security page for how proctoring data is stored and protected.
Human review of flagged attempts.
Proctoring flags — tab switches, webcam anomalies, activity breaks — are surfaced to administrators for human review, not automatically invalidated. An automated system cannot distinguish a candidate looking away from the screen from one consulting external material. The decision to invalidate an attempt is a human judgment call, made by a person with the context to assess it.
Anti-cheat design in question packs.
Question packs support randomisation of question order and, for MCQ, answer option order. Time limits are enforced server-side. Copy-paste is restricted in the coding environment. These measures reduce the value of answer-sharing between candidates sitting the same assessment drive.
Answer-key protection.
Correct answers and scoring rationale are never exposed to candidates — not in the assessment UI, not in share links, and not in any API response reachable by a candidate session. Access to answer keys is gated to administrator and reviewer roles only. This is enforced at the server layer, not by UI-level hiding.
AI-assisted scoring
How AI scoring works — and what it does not do.
For open-response and coding questions, AssessIQ uses AI-assisted scoring calibrated to the same rubric a human reviewer would apply. The scoring model is given the question, the rubric anchors for each band, and the candidate response — and returns a band score plus a rationale explaining which anchor the response matches and why.
Scoring is triggered by an administrator, not automatically.
AI scoring runs on administrator action — it does not fire automatically when a candidate submits. This is a deliberate design constraint: the administrator decides when results are ready to review, not an automated pipeline.
Every score carries a rationale.
The band score alone is not the deliverable. Every AI-scored response includes a rationale: what the candidate demonstrated, which anchor it matched, and why the score is not higher or lower. This makes the score reviewable by a human administrator and challengeable if a candidate disputes it.
Reviewers can override.
Administrators and assigned reviewers can inspect the AI-assigned score and rationale and override it. The final score on a candidate record reflects human judgment if a reviewer has intervened. AI scoring is an efficiency tool, not a black-box decision system.
Questions about assessment design?
If you have specific requirements — custom rubric design, fairness review of an existing pack, or questions about how scoring works for your use case — reach out. We will give you a straight answer about what AssessIQ can and cannot do for your context.
See how it applies to your context