If you are struggling right now, support is free and available 24/7. Find a helpline in your country

How we score, in full

Published so it can be checked. If our arithmetic disagrees with a published scoring manual, we want to hear about it.

Where the questions come from

Every instrument is transcribed from its published source (the original validation paper, the author's own distribution, or a government registry) into a structured record holding the items, the response scale, the subscale each item belongs to, the reverse-keyed items, and the normative cut-offs. Those records are regenerated from source on every deployment, so what you answer is what the instrument's authors wrote.

Reverse-keyed items

Well-built questionnaires deliberately word some items backwards, so that agreeing indicatesless of the trait rather than more. It is a check against people answering on autopilot.

Those items must be flipped before summing. On a scale running from minimum to maximum, a reverse item scores as min + max − response. On a 1–5 scale, an answer of 1 counts as 5. Skipping this step does not produce a slightly-off score. It produces a meaningless one, and it is among the most common errors on free testing sites.

Subscales are compared by average, never by total

When an instrument measures several things at once, comparing the raw totals is invalid unless every subscale has the same number of items, because otherwise the longest subscale wins by construction, regardless of how anyone answered.

We rank subscales by their mean item score, and report each one's position within its own possible range so that scales of different lengths stay comparable.

Three ways a result is interpreted

Which one applies is inferred from the instrument's own normative table, not assumed.

  • Single total. One score read against published cut-offs, such as the PHQ-9's 0–27 range, for instance.
  • Numeric profile. Several subscales, each read against the same table. The Big Five works this way: each trait is banded independently.
  • Categorical profile. Subscales combine into a named category, such as the four adult attachment styles.

The distinction is drawn by comparing the highest cut-off in the table against the maximum possible total and the maximum possible subscale score. If a table tops out at 20 on a hundred-point instrument, it is describing one subscale, not the whole thing.

We would rather show nothing than guess

If a score falls outside every published band and cannot be resolved to within a couple of points of one, no severity label is shown. A wrong label on a depression screener is worse than no label, so the breakdown is presented without one rather than rounding you into a category you are not in.

Risk detection

Some items ask directly about self-harm. The PHQ-9's ninth question is the clearest example, and endorsing it produces a total of only 1, a score that reads as "minimal" on every published cut-off table.

So risk is assessed separately from the score, before any result is displayed. Any endorsement above the scale's floor on such an item routes to crisis support first. This behaviour is covered by automated tests that block deployment if they fail.

What we do not do

  • No percentile claims we cannot support. Comparing you to a norming sample requires that sample's demographics. Where we do not have them, we report your position within the scale's range and say so, rather than inventing a percentile.
  • No adjustment of published cut-offs. Bands come from the validation research as published. We do not tune them to make results feel better or worse.
  • No inferences across instruments. We do not combine results from different questionnaires into a composite profile. That would need validation nobody has done.

Known limitations

Screening questionnaires measure how you have felt over a stated window, and they are sensitive to mood on the day. Several instruments here were validated on narrow populations (frequently Western, university-educated samples) and their cut-offs may not transfer cleanly to everyone. Where an instrument was designed to be administered by a clinician rather than self-completed, we label it as such.

Every result page names the sources behind it. Corrections are welcome.

Who wrote this page

Psychometrics Today Team, Editorial team. Last reviewed 30 May 2026.

The assessments on this site were developed and validated by the researchers named in each page’s sources. Psychometrics Today administers them and explains what results mean, in plain language, with every factual claim traced to the study it comes from.

Reviewed for accuracy against the primary research it cites. Not reviewed by a clinician, and not a substitute for assessment by one.