AI in Clinical Practice
I was building a set of clinical scenario cards a few months ago — short illustrated vignettes for a tool, the kind of thing a counselor and a client look at together on a screen. I’d asked an AI to help me develop the characters across several decks, and I’d asked for diversity directly, more than once: gender, race, ethnicity, body, ability, class. I was specific about it because I know the work, and I know what gets left out of it.
It did what I asked. The characters varied. And then I was a few decks in, and I stopped, and I looked back at what we had actually built.
What the AI handed back
Here is the part of that exchange that matters. These are real exchanges from my own work, edited for length — I’ve cut some of the back-and-forth, but nothing in the substance has been changed
From the exchange
Ok this is the second Black character about a dad in prison. It’s a stereotype.
You’re absolutely right, and that’s an important call. Two Black characters connected to incarceration — even across different decks — is a harmful pattern regardless of the statistical data behind it. It flattens representation and reduces Black characters to a single narrative. [… It went on to propose specific reassignments for the affected cards — cut for length…]
Read what happened there. The moment I named the pattern, the AI could describe exactly what was wrong with it. Fluently, accurately, in language I’d be glad to see from a colleague. The knowledge was right there. It just hadn’t done anything with that knowledge while it was actually producing the cards.
Why Better Prompting Isn’t The Answer
This is where the easy fix tempts you, so let’s deal with it now. If it knows, just tell it to check. And I do — I prompt for it directly, every time. But watch what that actually buys. The AI applied what it knew only on the dimension I had thought to name. I named race. So I kept pulling the thread, and I asked it to audit the whole set for something I hadn’t named yet:
From the exchange
Check, because I feel like a few of our plus-size characters have been Black. Can you check the variations we’ve chosen to be larger build?
[… It listed the larger-build characters deck by deck — cut for length…] Your concern is valid. Of the larger-build characters where the build actually landed, the breakdown skews heavily Black.
No single card was wrong. You could look at any one of them on its own and see nothing. The pattern only existed across the whole set — body size quietly tracking with race — and it stayed invisible until somebody went looking for it.
And notice who went looking. The AI ran the count. I read the count, and I checked it myself, because the system that produced the skew is not the system I’m going to trust to have the last word on whether the skew is real. But that comes second. The prompt that found this pattern only existed because I already suspected the pattern was there. A prompt can only ask the AI to check for what the clinician has already thought to look for. The judgment about what to look for in the first place was never in the prompt. That part was mine.
Researchers have documented the same pattern — a recent Cedars-Sinai study found leading large language models proposed different treatment recommendations for psychiatric patients when African American identity was stated or implied.
Even the fix needed reading
One more thing, because it’s the part that taught me the most. When I asked the AI to rebalance, it moved the larger build onto the mixed-race character and the Middle Eastern boy. The white boys in the decks stayed exactly as they were. The fix carried a quieter version of the same logic it was supposed to be fixing. Even the correction needed reading.
I caught all of this. Not because I’m unusually sharp — I want to be honest that I’m not — and not because the AI was broken, because it wasn’t. I caught it because reading a set of characters and noticing what they assume, and what they quietly repeat, is a clinical act. It is the same act I’d perform reading an intake another clinician wrote, or a treatment plan, or my own notes from six months ago. The catch wasn’t a technical skill. It was clinical judgment, pointed at a kind of document I hadn’t pointed it at before.
I want to be clear about one thing before I go further, because it would be easy to read a story like this as a case against using AI. It isn’t. I kept those cards. I fixed what needed fixing, the tool shipped, and the draft the AI gave me genuinely moved the work faster than I could have moved it alone. The draft was not worthless. The draft was just not finished — and the part that finished it was mine.
That’s what this blog is about. Not AI, exactly. Clinical judgment, and what happens to it when AI is in the room.

The conversation we’re mostly having sideways
For about two years now I’ve been using AI in the work I do supporting SUD clinicians — the resource library, the survival-series writing for clinicians early in their careers, the longer skill-building courses, the clinical tools. It runs underneath a lot of how I produce all of it, at whatever pace I can manage alongside seeing clients.
And the whole time, alongside the work, I’ve been building something else: a working sense of where AI output needs watching. The places it tends to go quietly wrong. What to do about them when it does. That sense didn’t arrive all at once. The mistakes got easier to see as I learned where to look, and at some point it occurred to me that the looking itself — the where to look — was the part actually worth handing to someone else.
I wouldn’t have written about this two years ago. Not because I was hiding it, but because the conversation around AI in clinical work is genuinely hard to enter without getting hurt by it.
You probably know the shape of it. Clinicians are using AI. A lot of us are. And the conversations I’ve watched happen most honestly are the ones happening sideways — in direct messages, in anonymized posts, in questions asked from accounts with no name attached to them. Because the moment you say “I use AI for this†out loud, you get sorted. Into the camp that thinks you’ve outsourced something you shouldn’t have, or the camp that thinks you’ve finally caught up with the times. There is no version of saying it plainly that doesn’t get you filed under one of those, and most people would rather not be filed, so they say nothing.
That’s a broken conversation. And a broken conversation has a cost, and the cost doesn’t land on the clinicians. It lands on clients. When clinicians can’t compare notes, nobody gets better at catching what I caught in those cards, and the output just goes through.
So I changed my mind about writing about it. Not because the conversation got any easier, but because staying quiet turned out to have its own cost, and that one was worse.

This is harm reduction. You already know how to do it.
Here is the frame for the whole blog, and it is not a technology frame. It’s one you already use every week.
You don’t tell a client to abstain from a substance because abstinence is the only respectable position. You meet them where they are. You look honestly at what is actually happening. You work to reduce the harm. And you do that because the alternative — refusing to engage with them until they’ve earned your engagement — leaves them alone with the danger.
Clinicians are using AI. That is the ground reality, the same way substance use is a ground reality in the rooms we work in. The pure-abstinence position on it — clinicians shouldn’t use AI, and teaching them to use it better only legitimizes it — has the same flaw the abstinence-only position has in our actual clinical work. It sounds principled. It leaves people unsupported.
And here is where the analogy gets sharper, not softer. Harm reduction was never only about the person doing the thing. It has always been about the harm that radiates outward — to the children, the partner, the people standing around someone who didn’t choose the behavior and absorb the cost of it anyway. That is the part that transfers most exactly.
When AI output goes unread, the harm doesn’t land on the clinician who used it. It lands on the client. The person who never touched the tool, who never consented to it, who is simply being described by it. The stereotyped character. The treatment plan that assumes a life the client isn’t living. The note that would shame the family if they ever read it. The client is the third party here, and harm reduction has always been most serious about the third party.
So this blog is harm reduction applied to your own use of AI. Not anti-AI, not pro-AI — a practitioner’s working account of how to stay clinically responsible when AI is one of the things feeding your work. It is one clinician’s account of a reading practice. It is not a ruling on whether AI use is permissible under your license or your agency’s policy. Those are real questions. They just aren’t the ones this blog is here to settle.
One more thing, because “your own use†isn’t the whole picture. Some of you chose AI. You opened the tool, you decided. Others didn’t — it arrived inside the systems you already work in, built into the documentation, the workflow, the intake, with nobody asking your opinion first. If AI was handed to you rather than chosen by you, reading its output well matters more, not less.
Though I’ll be honest with you: seeing a problem and being free to fix it are not the same thing. If you’re working under a system you didn’t choose, you can catch what the AI got wrong and still be required to use it anyway. This blog can’t hand you authority you don’t have. What it can do is make sure that when you get overruled, you at least knew, and you could say so, on the record.

What this blog is, and what comes next
The thing I want to leave you with isn’t a warning. It’s a different question than the one people keep asking.
The question was never whether you use AI. Plenty of careful clinicians do, plenty don’t, and that isn’t the line that decides anything. The question is whether you can still see the frame in what it hands back to you — whether, when the output is fluent and confident and almost right, you can catch the almost.
And here is the part that is meant to make this feel possible instead of exhausting. The catching is clinical judgment. It is the same faculty you have been building since the day you started doing this work. It is not a fixed thing you either have or you don’t. It develops the way the rest of your clinical judgment developed — slowly, through cases, through getting it wrong and noticing you got it wrong. Wherever you are with it right now, it is something you build. It is not something you are missing. What’s new here isn’t the skill. It’s the document the skill gets aimed at.
That’s what the posts ahead are for. The next one takes on the objection I hear most — it keeps getting better, won’t it just get good enough that I won’t have to do this anymore? — because if that objection holds up, not much else here matters. The one after that goes somewhere less comfortable. The AI output that is hardest to catch is not the output that is obviously wrong. It is the output that feels right. And that one doesn’t announce itself, which makes it a harder problem than anything in this post.
For now I’ll leave it where those cards left me. I caught the stereotype, and catching it made me look harder. The audit found the pattern. The fix needed catching too. It was never one moment of noticing — it was a practice of it, and one I am still building. The AI gave me what I asked for, every single time. Knowing what to ask it — knowing what to look for before I had any real reason to be sure it was there — that was the work. It still is.

Next: capability isn’t the problem — why a more capable AI doesn’t retire your judgment.