psisula
Blog

Making assessment scales part of the routine

Turning PHQ-9 / GAD-7 into a habit, letting scoring run itself, and reading the trend that comes out of it.

The psisula team3 min read
assessmentsoutcomes

Administering a scale and using the data from it are different things. In most practices a PHQ-9 gets filled in at intake, the score goes in the file, and the file never opens again. A single score is a snapshot. The clinical value is in the slope that only appears once there’s a second and a third measurement.

Rhythm first, catalogue second

The common mistake is to start by shopping for a large scale library. What actually changes your practice isn’t how many instruments you can reach — it’s how many clients you measure on a regular interval. Two scales applied consistently across eight sessions tell you far more than fifteen scales applied once.

A cadence that works to start with:

  • Intake: one primary scale matched to the presenting problem (PHQ-9 for depression, GAD-7 for anxiety), plus a screener if warranted.
  • Every 4 sessions: the same primary scale again. Holding the interval steady is what makes the slope readable.
  • Before termination: a final measurement, compared against the baseline.

Delivery friction sets your measurement frequency

The real obstacle to measuring isn’t clinical, it’s logistical. Printing paper forms, spending ten minutes of session time, or saying “fill this in sometime” and losing track of it will wreck your interval fast.

When the client completes the scale in the portal, that step leaves the session entirely: the form goes out beforehand, the client fills it in on their own time, and the score is in the file before you walk into the room. In-session entry stays available too — that option shouldn’t close for clients without a phone or who need help working through the items.

What automatic scoring removes

Hand-scoring produces two kinds of error: arithmetic slips, and missed reverse-coded items. Both are silent — a wrong score looks exactly like a right one and gets buried in the trend.

Automatic scoring removes that, and it also makes subscales usable. The depression/anxiety/stress split in DASS-21, or the nine dimensions of SCL-90-R, tend to get skipped in practice when they have to be computed by hand. Arriving already computed is what makes them clinically usable.

A threshold flag is a signal, not a decision

When a clinically significant threshold is crossed, a rule-based flag appears on the results screen — for example, a non-zero response on the PHQ-9 suicidal-ideation item. This is not a risk assessment and it does not substitute for one. The only thing it does is make an item visible that could otherwise be skimmed past on a busy day. The judgment stays yours.

Reading the trend

  • Don’t interpret a single measurement. One bad score doesn’t mean the work is going badly; it may just have been a bad day.
  • Hold the interval steady. Scores taken at irregular intervals produce noise that looks like a slope.
  • Read it alongside session density. A rise during a three-week gap may show that treatment wasn’t applied, not that it isn’t working.
  • A plateau is information too. The point where improvement stops is usually the point to revisit the formulation.

Because score history is kept per client, you can export that curve as a PDF for a supervision session or a referral letter. The client sees their own curve in the portal as well — tying a vague sense of “this is going well” to something concrete supports the work in its own right.

That’s the payoff of making measurement routine: instead of arguing about whether treatment is working, you can show it.