How we build the test

The research our items rest on, how we select and monitor them, how we check reliability and validity, and how we turn a raw score into a percentile and an IQ-scale number.

Last updated 30 June 2026 · written and reviewed by Marcus Hale

Most online IQ tests hand you a number without ever explaining where the questions came from or how an answer became a score. We work the other way around. This page documents our process end to end, from the research the items rest on to the standard error we report with every result, so you can judge the test rather than trust it on faith.

Why our approach is different

The honest difference is not a secret algorithm; it is that we use validated items instead of inventing our own, and we publish the maths instead of hiding it. A test is only as good as two things: where its questions come from, and how it turns answers into a number. Both are written out below in plain terms. Where the format imposes limits, we say so rather than papering over them, and where a figure carries a margin of error, we report the margin. Our scoring formulas live in full on the how we score page, and the standards we hold the whole operation to are set out in our editorial policy.

Stage 1: Choosing an evidence-based foundation

The test begins with a research-backed item pool, not a pile of puzzles assembled over a weekend. Our items are modeled on the International Cognitive Ability Resource (ICAR), a public-domain set developed by David Condon and William Revelle and documented in Intelligence (2014, vol. 43, pp. 52–64). We chose an established pool for one plain reason: ICAR items have been administered to large samples and correlated against full-length instruments, so they carry real measurement evidence rather than a designer's hunch. The structure also tracks the broad ability factors that decades of research, from Spearman's g to the Cattell–Horn–Carroll framework, have repeatedly recovered. Building on that work means our test inherits validation we could never replicate from scratch.

Stage 2: Selecting and balancing items

From the ICAR pool we select items across four families chosen to sample reasoning broadly rather than test one narrow skill: matrix reasoning, verbal reasoning, numerical and letter series, and three-dimensional rotation. Within each family we calibrate difficulty so the test contains easy, moderate, and hard items, which is what lets it tell apart people at different ability levels rather than bunching everyone together. We weight the set toward matrix items because they are the strongest single signal of fluid reasoning. We also keep wording as clear and as culturally neutral as the format allows, knowing that an ambiguous or knowledge-dependent item measures the wrong thing. The result is a profile across four domains, not a single trick that is easy to game.

Stage 3: Monitoring how items perform

Selecting good items is the start, not the end; we watch how they actually behave on real responses. For each item we look at its difficulty (the share of people who answer it correctly) and its discrimination, meaning how well performance on that single item tracks performance on the test as a whole. An item whose answers are essentially unrelated to the total score (a low item-total correlation) carries little information, and we flag it for review or removal. As a working rule we keep items whose correlation with the relevant scale is adequate, conventionally around 0.30 or better, and we retire items that drift below that or whose answer patterns look anomalous. This monitoring is continuous, because an item that performs well in one period can degrade as the audience or the wider context shifts.

Stage 4: Reliability

Reliability is how consistently a test measures, and it sets a hard ceiling on how much trust any single score deserves. We assess internal consistency, the degree to which items that are meant to measure the same thing agree with one another, using coefficients in the Cronbach's alpha family; the conventional bar is α ≥ 0.70 for acceptable and ≥ 0.80 for good, and that is what we aim for. Where the data allow, we also consider test–retest stability, how closely a person's results line up across two sittings, while remembering that the practice effect inflates a second attempt. We are candid that a short, unproctored online test cannot match the reliability of a supervised two-hour battery. Lower reliability means a larger margin of error, and we carry that fact straight into how we report your score rather than hiding it.

Stage 5: Validity

Validity asks the harder question: does the test measure what it claims to? Content validity is about coverage, whether the four item families between them sample reasoning broadly rather than leaning on one ability; our domain spread is the deliberate answer to that. Construct validity is about behavior, whether the scores act the way theory predicts, for example that the matrix items load most heavily on fluid reasoning. Because our items come from a validated public-domain pool, the test inherits a body of evidence linking ICAR performance to established cognitive batteries, with reported correlations strong enough to make it a dependable screen. What that evidence does not do, and what we never claim, is make an online screen equivalent to a clinical instrument.

Stage 6: Norming

A raw count of correct answers means nothing until it is compared with how other people perform, so the final stage places your score in context. We build normative distributions from accumulated responses and convert raw scores into interpretable standard scores: percentiles, z-scores, and the deviation-IQ scale with a mean of 100 and a standard deviation of 15. Where it is meaningful, we account for the fact that norms can differ across groups. Importantly, norms are not set once and forgotten; measured scores drift over generations (the Flynn effect), so a frozen reference sample slowly inflates everyone's percentile. We therefore recompute norms as data accumulates. The full conversion, including the formula and a worked example, is on the how we score page.

Transparency and limits

We would rather under-promise than oversell. We report the standard error of measurement, so a result is read as a band rather than a pinpoint, and we never assert equivalence to professional clinical instruments. The online format carries limits we do not control: we cannot guarantee the testing conditions, rule out distractions, or verify that someone answered without help. Those limits are stated on the result screen and spelled out in full on our disclaimer & responsible use page, which also explains how to read a score without over-reading it.

Continuous improvement

A test is never truly finished. We keep collecting responses, recomputing norms, and reviewing the items that perform weakest, and we revisit the test when the underlying research moves or our periodic review flags something, on the schedule set out in our editorial policy. The methodology and scoring are owned by Marcus Hale, who leads test development here. If you spot something you think is wrong, in an item or in this account of how we work, tell us through contact; corrections are genuinely welcome.