A depression test can flag symptom patterns and suggest a severity range, but it cannot reveal the underlying cause, confirm a diagnosis, measure daily functioning, or replace the clinical judgment a licensed therapist provides through a full conversation about your history and circumstances.
What if the number at the end of your depression test can't actually tell you what's wrong? Screening tools are built to flag patterns, not explain the person behind them. Here's what that score really means, and why the conversation after it matters more than the total itself.
What is depression screening?
Depression screening is a short, standardized set of questions built to flag the possibility of depression, not to confirm it. Every person answers the same wording in the same order, which is what makes a screening instrument useful. That consistency lets a clinic compare one person’s answers against a general pattern instead of relying on an open-ended conversation that might miss something. A screening tool is fast by design, often taking only a few minutes, and that speed is a tradeoff, not a flaw.
A score from one of these tools is not a diagnosis. Diagnosis is a judgment made by a qualified professional, who weighs the questionnaire alongside your history, your current life circumstances, and a direct conversation with you. The distinction between depression screening vs diagnosis matters because a score is only one input into that judgment, not the conclusion itself. Two people can land on the same number and be in very different places, which is part of why a clinician looks past the score itself.
Most people encounter depression screening without asking for it. It shows up as a routine part of a primary care visit, a student health appointment, a prenatal or postpartum checkup, or a general medical intake. That is by design: making it routine catches people who would never bring up their mood on their own. You do not need a reason to be handed one.
A screening score also describes a narrow window, usually how you have felt over the past two weeks, not a verdict on your life or your character. A rough two weeks can produce a high score even if that stretch is not representative of how you usually feel. If you want more background on what depression involves beyond a questionnaire, this overview of depression is a useful starting point.
How does the PHQ-9 work, item by item?
Depression screening tests work by turning a list of symptoms into a number. The PHQ-9 test for depression asks about nine specific experiences over the past two weeks, and you rate how often each one has shown up. Those answers get added together into a single score, and that score sorts into a severity band. The structure is simple by design, which is part of why it shows up in so many primary care offices and therapy intakes.
The nine items and what each one is asking about
Each of the nine questions maps to one of the symptom criteria used to diagnose major depressive disorder. The items ask about low interest or pleasure in activities, feeling down or hopeless, sleep problems, low energy, appetite changes, poor self-image, trouble concentrating, noticeable slowness or restlessness, and thoughts of death or self-harm. Nothing about the wording is abstract. Each question asks about a specific, observable pattern from the last two weeks, not a general sense of how life is going.
How responses become a score, and what the severity bands mean
For each item, you choose how often the symptom occurred: not at all, several days, more than half the days, or nearly every day. Each answer carries a fixed point value, from 0 for “not at all” up to 3 for “nearly every day.” Add up all nine items and you get a total score between 0 and 27.
The original validation of the PHQ-9 established the severity bands still used today:
- 0 to 4: minimal
- 5 to 9: mild
- 10 to 14: moderate
- 15 to 19: moderately severe
- 20 to 27: severe
These PHQ-9 severity bands describe symptom load over the prior two weeks. They say nothing about urgency, how long symptoms will last, or what kind of support fits a given person. A later meta-analysis of 18 validation studies found that cut-off scores between 8 and 11 all show acceptable diagnostic properties, which is part of why a single fixed threshold gets treated as a guide rather than a hard line.
The two items that are not part of the total
Two parts of the PHQ-9 sit outside the main scoring. The first is a tenth question, not counted toward the total, that asks how much these symptoms have interfered with your work, home responsibilities, and relationships. It is scored and reviewed on its own because two people can land on the same total score with very different levels of day-to-day disruption.
The second is item 9, which asks about thoughts of death or of hurting yourself. Regardless of the total score, a positive answer here gets flagged and looked at independently. If thoughts like these are something you are experiencing right now, help is available and does not require an appointment.
What other depression screening tools exist?
A depression questionnaire is rarely the only instrument in the room. Clinics and primary care offices pull from a range of tools depending on how much time they have, who the patient is, and what else might be going on. Recognizing the format helps you understand why you were handed a particular one.
Ultra-brief and long-form self-report instruments
The PHQ-2 asks about only two things: low mood and loss of interest, over the past two weeks. It exists to answer one question fast, whether a longer questionnaire is worth giving at all. A comparison of case-finding tools found that a two-question instrument covering depressed mood and anhedonia performed about as well as six longer, previously validated instruments, with sensitivity near the top of the range and specificity in line with the longer tools. That is the tradeoff built into the PHQ-2: speed first, detail later if needed.
On the longer end sits the Beck Depression Inventory, a self-report scale that leans heavily on cognitive and attitudinal symptoms: guilt, hopelessness, self-criticism, alongside the physical ones. It takes more time to complete than a two-item pre-screen and asks the person to reflect on how they see themselves, not just how they feel physically.
Instruments built for specific populations
Standard wording does not always fit the person answering it. The Edinburgh Postnatal Depression Scale was built for the perinatal period, when fatigue and appetite changes are expected parts of new parenthood rather than symptoms on their own. The Geriatric Depression Scale drops physical-symptom items that overlap with normal aging or chronic illness in older adults, focusing instead on mood and outlook. Both exist because a generic checklist can misread ordinary life stage changes as depression, or miss depression hiding behind them.
Self-report versus clinician-rated scales, and what are the best depression tests available
There is no single best depression test. The right one depends on setting, population, and how the answers will be used. Self-report tools like the PHQ-2 and GAD-7 are filled out by the person being screened, while clinician-rated scales such as the Hamilton Rating Scale for Depression are scored by an interviewer asking structured questions and judging the response. That shift changes what the score represents: one reflects self-perception, the other reflects a trained observer’s read on presentation. The GAD-7 itself measures anxiety, not depression, and often rides alongside a depression screener because the two conditions frequently overlap.
Free online quizzes complicate this picture further. Some reproduce a validated instrument item for item and score it the same way a clinic would. Others are informal content with no research behind the questions or the cutoffs, offering a number with nothing to support what it means.
How accurate are depression screening tools?
How accurate are depression tests? It depends on what you mean by accurate. A screening tool can be very good at catching people who have depression and still produce a lot of false alarms, because those are two different measurements, not one. Understanding both is the difference between reading a score as a verdict and reading it as a probability.
Sensitivity, specificity, and why cutoffs are a tradeoff
Sensitivity is how well a test catches people who actually have depression. Specificity is how well it clears people who do not. A meta-analysis of 17 PHQ-9 validation studies found sensitivity of 0.80 and specificity of 0.92 against a diagnosis of major depressive disorder, meaning the PHQ-9 correctly flagged about 80 out of 100 people who had depression and correctly cleared about 92 out of 100 who did not. Those two numbers move against each other. A meta-analysis of 18 PHQ-9 studies across different cutoff scores found specificity ranging from 0.73 at a cutoff of 7 up to 0.96 at a cutoff of 15, with sensitivity shifting the opposite direction as the cutoff rises. Raising the cutoff catches fewer false positives but misses more real cases. Lowering it does the reverse. There is no cutoff that eliminates the tradeoff, only one that shifts where the errors land.
Why the same score means different things in different settings
The question you actually want answered is different from sensitivity or specificity: given that you scored above the cutoff, how likely is it that you meet criteria for depression? That number is called positive predictive value, and it depends heavily on how common depression is in the group being tested. The same instrument with the same cutoff behaves differently in a general waiting room than in a specialty mental health clinic. In a low-prevalence setting, even a specific test generates enough false positives to outnumber the true positives, simply because so few people in the room have the condition to begin with. In a high-prevalence setting, like a clinic where most patients already show symptoms, that same positive score is far more likely to reflect a real case.
What a false positive and a false negative actually feel like
A false positive does not mean you were dishonest or exaggerating your answers. It means a short questionnaire cannot always tell depression apart from grief, burnout, sleep debt, thyroid problems, or the aftermath of a genuinely hard month. A false negative means the questionnaire missed something real, which is part of why these instruments are validated against a reference standard, usually a structured diagnostic interview, that is itself an imperfect measure. A score is a starting point for a conversation, not a diagnosis on its own.
What depression tests cannot tell you
What depression tests can and cannot tell you
A depression test can tell you whether your answers add up to a number that falls inside a range a clinician would want to look at more closely. What it cannot tell you is why the number landed where it did. The same total can sit on top of a recent loss, a chronic illness that has worn someone down for years, a relationship that is falling apart, or nothing the person can point to at all. Depression screening limitations start here: the score has no way to hold context, only content.
A score also cannot separate depression from other things that produce similar answers on paper. Anxiety, trauma responses, ADHD-related exhaustion, hypothyroidism, anemia, chronic pain, and disrupted sleep can all push someone toward the same checked boxes. The questionnaire does not know which one it is looking at. If you want a fuller picture of depression as a condition with many possible underlying causes, that distinction lives in a broader resource on depression, not in the test itself.
The gap between the answer and the truth
A test cannot register masking, which is the practice of answering at the level you believe is acceptable rather than the level that is true. Because depression exists internally rather than visibly, someone can minimize their symptoms on paper even when sitting across from a clinician, and the questionnaire has no mechanism to detect that gap between the answers given and the experience lived.
Dr. Tina Fornwald describes this in terms of how invisible pain gets minimized: “Unfortunately, we cannot see the profound trauma that we’re dealing with, so we may underestimate how much it’s impacting us. But if you had your arm that was severed, you could see that pain and everyone around you could see that and they would support that. Sometimes because we just don’t want other people to know what’s going on, we may come up with, “I’m okay.””
Shame is one force behind that underreporting. Kristen McLoud puts it this way: “It’s something that we’re ashamed of a lot of the time. How could I have that thought about myself? Or how could I have that thought about anything random going on? And it’s just something that the shame can lead to us backing away from getting help.”
What the list leaves out
Most screening tools ask about a fixed set of symptoms, and plenty of real presentations fall outside that set. Irritability, numbness, physical heaviness, a loss of recognizing yourself, or a flatness that never quite reaches sadness can all be part of depression without ever getting their own question. A test also cannot capture function. Two people can land on the exact same total while one is still working, eating, and talking to people, and the other has stopped doing all three.
