A free depression test measures symptom severity using validated tools like the PHQ-9, but it cannot diagnose depression, rule out conditions such as bipolar disorder or grief, or replace a clinical interview with a licensed therapist who evaluates your full history and context.
What if a positive result on a free depression test is right only half the time? That's not a flaw, it's math, since accuracy shifts depending on who's taking it. Here's what your score can flag, what it can't, and what actually happens next.
What a free depression test is, and what it measures
A free online depression test is almost always a self-report questionnaire. You read a short list of statements and rate how often each one has applied to you, rather than being observed or examined by anyone. There is no lab work involved and no one watching how you answer. The format is simple by design: a handful of questions you can complete in a few minutes on your phone or computer.
Most short depression tests ask about the same core set of experiences. Depression symptoms typically include persistent low mood, loss of interest in things you used to enjoy, changes in sleep or appetite, low energy, trouble concentrating, feelings of worthlessness, and thoughts of death or self-harm. A screening questionnaire is built to capture that same list in a structured way, usually asking you to rate each item over the last one or two weeks. That recall window matters because it anchors the test to a recent stretch of time rather than asking you to summarize your whole life.
A screen like this measures symptom burden at one point in time. That is a different task from identifying a specific condition or explaining what caused it. Screening exists to be a fast, low-cost way to flag who might benefit from a closer look, and it is built on purpose to be over-inclusive rather than to miss people who need support.
Some sites ask for an email address before showing your results, while others require none. Neither version is more or less valid: the difference is how the site collects your data, not how it measures your symptoms. You will also find sites that combine a depression and anxiety screen into a single form, since the two symptom sets overlap and often occur together. If you want to read more about the condition these tools are trying to screen for, depression is covered in more depth elsewhere on ReachLink.
The eight depression screening instruments, compared
Eight instruments show up most often when people search for a depression test, and they were not built for the same person or the same purpose. Some are meant to be filled out alone in a waiting room. Others exist only as part of a clinician-led interview. Knowing which one you have taken tells you what the result can and cannot say about you.
Self-administered questionnaires you can take on your own
The PHQ-9 is the instrument you are most likely to encounter, including on most general-interest screening sites. It has nine items, each mapped directly to a diagnostic symptom of depression, and it was built and validated as a self-administered screen in primary care and obstetrics-gynecology settings. It takes only a few minutes to complete and is free to use.
The PHQ-2 is the short version: just the first two PHQ-9 items, asking about the frequency of low mood and loss of interest over the past two weeks. A 2003 validation study in Medical Care reported 83% sensitivity and 92% specificity for major depression at a cutoff score above 3, and identified a score of 3 as the optimal screening cutpoint. It was designed as a first-pass filter rather than a standalone result, meant to flag who should take a longer instrument next, not to stand in for one. It is free and self-administered, and takes under a minute.
The BDI-II (Beck Depression Inventory) is a longer self-report questionnaire, 21 items, with more weight given to cognitive and attitudinal symptoms such as self-criticism, guilt, and pessimism than the PHQ-9 carries. It takes longer to complete, usually five to ten minutes, and unlike the PHQ-9 it is not free to use in most settings.
The Zung Self-Rating Depression Scale is an older self-report tool that mixes affective, physiological, and psychological items into 20 questions. It still turns up on free test sites, though it predates the PHQ family and is used less often in primary care today.
Scales built for specific populations
The Edinburgh Postnatal Depression Scale (EPDS) was built specifically for pregnancy and the postpartum period, because standard scales tend to confuse normal pregnancy symptoms, like fatigue or appetite changes, with depressive ones. It is a 10-item, self-administered questionnaire and it is free.
The Geriatric Depression Scale (GDS) was written for older adults. Its wording deliberately avoids somatic items, physical complaints that in an older adult are more likely to reflect a medical condition or normal aging than a mood disorder. It is self-administered, free, and comes in a 30-item and a shorter 15-item version.
Clinician-rated scales you will not find as a free online quiz
The Hamilton Depression Rating Scale (HAM-D) is administered by a clinician, not filled out alone, and it has historically served as the outcome measure in depression treatment research. The Montgomery-Asberg Depression Rating Scale (MADRS) is also clinician-administered and is weighted toward detecting change over time, which is why it shows up in trials tracking whether symptoms improved rather than in general screening. Neither appears as a free online test, because both require a trained rater asking questions and scoring responses in real time.
How depression test scores are calculated and what the ranges mean
A short depression test works by turning each answer into a number. Most self-report screens ask how often a symptom showed up over a set number of days, then assign a value to each response based on frequency. Those values are added together into a single total. That total is the number you see at the end of most free screening tools.
Why the same number means the same thing everywhere
Severity bands exist so that a score carries a consistent meaning no matter where it is used. A 2001 validation study in the Journal of General Internal Medicine found that PHQ-9 scores of 5, 10, 15, and 20 mark the boundaries for mild, moderate, moderately severe, and severe depression. These bands describe symptom load, not a diagnosis and not a fixed identity. A score of 12 places you in a range other people also land in, but it does not say what put you there.
A cutoff is the point at which a screen recommends a closer look. On the PHQ-9, a score of 10 or higher identified major depression with 88% sensitivity and 88% specificity when checked against structured clinical interviews in that same study. That threshold was chosen to balance catching real cases against flagging people who do not have the condition. Other settings sometimes use a different cutoff for the same tool.
Why the pattern behind the number matters
Two people can reach the identical total through completely different symptoms. One person’s score might come mostly from sleep and appetite changes, another’s from loss of interest and low energy. Functional impairment questions, often placed at the end, ask how much these symptoms interfere with work, home, and relationships, and they carry weight on their own, separate from the sum. Any answer indicating thoughts of death or self-harm is also handled apart from the total rather than folded into it.
If you are having thoughts of hurting yourself, help is available right now and it does not require an appointment.
A single score is a snapshot. Watching a number rise or fall across repeated administrations often tells you more than any one result on its own.
What a positive result actually means, worked through with real numbers
A score above the cutoff on any accurate depression test is not a verdict. It is a signal with a known error rate, and that error rate changes depending on who is taking the test. Understanding two statistics, sensitivity and specificity, along with a third idea called predictive value, explains why the same number can mean something different for you than it did for a friend who scored the same way.
Sensitivity and specificity in plain language
Sensitivity is how often a test correctly flags people who actually have depression. Specificity is how often it correctly clears people who do not. A 1997 study in the Journal of General Internal Medicine comparing case-finding instruments in primary care found that a brief two-question screen had 96% sensitivity and 57% specificity against a structured diagnostic interview, while the 2001 PHQ-9 validation study found 88% sensitivity and 88% specificity at a cutoff score of 10. Neither number answers the question you actually care about, which is the reverse one: given that you scored above the cutoff, how likely is it that you are depressed. That reverse question has its own name, positive predictive value, and it depends heavily on how common depression is in the group being tested.
Why the same score means different things in different settings
Say a free screening questionnaire is given to 1,000 people in a general community setting, where depression is relatively uncommon. Using the PHQ-9 figures of 88% sensitivity and 88% specificity, and assuming roughly 100 of those 1,000 people actually have depression, the test correctly flags about 88 of them, while also flagging about 108 of the 900 people without depression as false positives. More of the flagged people are false positives than true ones, which means fewer than half of the positive results in that group reflect actual depression. Now imagine the same test run where depression is much more common, say 4 in 10 of the group. It correctly flags about 352 true cases against only about 72 false positives. These figures are illustrative rather than a description of any particular clinic, but they show how much the setting changes what a positive result means.
When a self-report score does not match the person
The practical takeaway is that a positive result is a reason to look further, not a conclusion that stands alone, and this is especially true for anyone in a low-prevalence group. Negative predictive value on these instruments tends to run high, so a low score is fairly reassuring for the specific symptoms asked about, though it says nothing about a condition the questionnaire never asked about. Scores can undershoot reality when someone minimizes their answers, describes distress through physical complaints instead of mood, comes from a background where low mood is not expressed openly, or answers the way they think they are supposed to. Scores can overshoot it when grief, acute stress, sleep deprivation, physical illness, or medication effects produce symptoms that look identical on paper to depression, because a questionnaire measures symptoms, not their source.
Charity Anderson, LPC, traces that last kind of reticence back further than any questionnaire. Speaking on the ReachLink podcast about where resistance to talking about mental health comes from, she put it this way: “the misconception of therapy is born in childhood. When we teach our children, what happens in my house stays in my house, and you don’t tell nobody what’s going on, you’re teaching your children that it’s not okay to talk to people, that it’s not okay to express yourself.” Someone raised on that message can answer every item on a screen honestly and still land low, because what they learned to say about their mood is not the same as the mood itself.
What happens between a positive screen and a diagnosis
An elevated score on a free screening questionnaire is a starting point, not an endpoint. What follows is a defined sequence, and knowing the steps helps you understand why the process takes longer than filling out one form.
From questionnaire to clinical interview
The first step is usually a repeat or a fuller version of the same self-report screen. A single high score during a genuinely bad week reads differently than a pattern that holds steady over time, so a second measurement helps separate the two. The step after that is a clinical interview, and this is where a form stops being enough. An interviewer may ask questions in an open, unstructured way, or may use a structured tool such as the SCID, which gives trained interviewers a standardized sequence of questions that branches based on the answers. The MINI works on a similar principle. Either way, the interview responds to what you say instead of moving down a preset list.
Checking symptoms against formal diagnostic criteria
The interview responses get checked against the criteria written into the DSM-5-TR, the manual clinicians use to define mental health conditions. Those criteria specify which symptoms count, how many need to be present, how long they need to have lasted, and whether they are severe enough to interfere with daily functioning. An accurate depression test can flag distress. It cannot apply that full set of thresholds on its own, which is the real reason the interview step exists.
