The Underrated Superhero

Resources
for Clinicians

It Recommended the Safe Thing.

AI Training Data Bias


I was drafting a note for a client with high acuity. Everything I put in front of the tool was cleansed of identifying information first, which isn’t optional and isn’t the point of this post, but I know someone is about to ask.

What I wanted was a note that did what notes have to do. Document the contact. Capture the acuity, because the acuity is fact. Show what I was doing about it, clinically and ethically, in a way that would hold up if anyone ever looked.

So, I wrote what we’d discussed. Hospitals were discussed, and here’s why that wouldn’t work. Groups were discussed and ruled out quickly due to client’s behavior. Higher level of care wasn’t currently an option. Here’s what the client and I landed on instead.

The reasons were all in the prompt, written by me. What I didn’t do was tell it what to do with them — how much of that belonged in the note, or that it should be framed as the rationale for the plan rather than just recounted as session content. It worked that out on its own, and that inference is exactly why I use these tools. It read what I meant, not only what I typed.

Then, in the same draft, it added a line saying I would encourage increasing the check-in to a full session given the acuity.

Not a suggestion to me. A line in the chart, in my voice, ready to sign.

Not a suggestion to me. A line in the chart, in my voice, ready to sign.

More contact was already what we’d built in the plan. The client and I worked out check-ins on the off weeks — case management, brief, documented. The client was afraid of being charged. It took a while to earn enough trust for the client to accept them at all.

It had all of that in the prompt, and the line it wrote would have converted the plan back into the thing the client couldn’t do. It reached for the standard shape anyway.

What I didn’t ask for

And notice those are the same move. The thing that took my session content, recognized my intent, and documented it as such is the thing that decided I must also want a recommendation. One matched what I wanted. One didn’t. It had no way to tell them apart, and they arrived in the same paragraph, in the same voice.

I understand the pull toward that standard shape, too, because I felt it. High acuity is hard to sit with. Everything in you wants to do more, because doing more is what relieves the feeling. Whether it helps is a separate question, and it’s the one you can’t answer from inside the pull. I had to keep coming back to what was actually in my control and what wasn’t, and that isn’t a thing you settle once. So, when the tool reached for more contact, it wasn’t reaching for something foreign to me. It reached for the thing I’d already had to talk myself out of. Not because I wanted less for the client — because more, in that shape, wasn’t something the client could use.

More contact is the safe recommendation. Not because it’s clinically stronger, but because it’s the one nobody second-guesses. If something goes wrong later, I increased contact is the line that survives review. That’s what makes it safe, and it’s why the suggestion doesn’t have to be dramatic to be the standard.

So, this is a small story, deliberately. It also isn’t a story about why you shouldn’t use these tools. I use them every working day and they save me hours. It’s about the one thing you have to keep doing anyway.

This is the sixth post in a series, and it stands on the ones before it. If standpoint, the echo, or recognition aren’t familiar terms, start at the beginning — that’s where they earn their meaning.

What actually happened there

I want to be careful about what I’m claiming, because I’m not saying I made the right call. I don’t get to know that, and neither does anyone else — that’s been true since the last post. Another clinician could look at the same picture, push harder toward more intensive care, defend it, and be called thorough for it.

Sometimes we do know what it costs. We report, and the family fractures. We hospitalize, and the person never tells a clinician the truth again. Sometimes that’s the call and it’s still the right one. Sometimes it’s fear, and the client absorbs it. The uncomfortable part is that the same language covers both, and one of them is never going to be questioned.

The same language covers both, and one of them is never going to be questioned.

What I can say is narrower. I gave it the grey, and it gave me back something with an answer in it.

By the grey I mean the distance between what the textbook says is best and what’s actually available. No one to refer to, when you know the referral is what’s indicated. An IOP that’s right for the client and that the client refuses. Harm reduction, in nearly any form. The least restrictive option, chosen and then sat with.

A note on “what the textbook says,” because for a lot of us the textbook isn’t a neutral authority. It’s a body of research that was built from and for a small subset of clients. If you work with justice-involved people, in harm reduction, or anywhere the funding didn’t follow, you already know the feeling of practicing in territory nobody wrote the study for. That’s not the exception to the grey. For a lot of us it is the grey, most of the time.

Working outside the best-case version isn’t the same as working outside evidence. You’re still accountable to something: to what this client will actually do, to what’s reachable, to whether it’s working, revised when it isn’t. The clinician sitting with the least restrictive option because the IOP was refused isn’t freelancing. They’re accountable to the person in front of them instead of to a plan written for a different one.

What came back in that note was the standard answer. The one that would look thorough to anyone reading over my shoulder.

Post two named what a lot of documentation AI products are actually oriented toward: getting paid, surviving audit, satisfying medical necessity. That’s the note’s language. This is the same pull reaching the clinical decision itself. Mine was the small version — one line in a note, easy to miss.

The same move shows up at scales you’d catch. A referral the client can’t use. A level of care that takes the work out of their hands. When the referral is all the client gets, our ethics code has a name for that, and the name is abandonment.

Those are the visible ones. Mine was a line in a note, and that’s the version that does the damage, because it can be easily missed. It doesn’t get argued about. It gets signed, and then it’s in the record, and the next one goes in the same way.

The defensible move and the right one come apart more often than we say out loud, and the tool only has a strong signal for one of them.

Last post I pushed back on it and it folded — abandoned a position it had good reason to hold, because I doubted it. This time it held. Both times, the thing that moved the output was something other than the person the output was about: my push-back then, the weight of the standard now. That’s why neither one is evidence about the client.

Why the standard answer is the one that shows up

The answer is the whole Foundation arc and I won’t rebuild it here: the output is rendered from an aggregate, and the aggregate leans toward whatever was written down most. Post three worked the version of this you already feel when you’re asking — push for a frame the literature is thin on, and the answer keeps sliding back to the dominant one. And when it does get written up, it tends to arrive defending itself. I wrote a PHI disclaimer into the third sentence of this post, and nobody made me, I just anticipated criticism.

Here’s the shape of the standard answer in documentation. The medical frame is written down as procedure. Indications, criteria, escalation pathways, what to do at what level of acuity, organized around the physician’s expertise because that’s what medicine documents. It’s thorough and it’s specific and it tells you what to do next.

Meeting a client where they are is written down as principle. So is client autonomy, so is the person’s right to decide what they’ll actually accept. All of it is real and evidence-based and taught — and almost none of it arrives as procedure, because it can’t. The workaround is always specific. Increase frequency given acuity has a procedural shape. Increase frequency in whatever form this particular person will actually accept, arrived at over weeks, for reasons that live entirely in their life does not. It gets assembled per person, which is exactly what keeps it out of the kind of text you can draw a recommendation from.

So the two frames aren’t competing on merit. They’re competing on form, and procedure wins that every time something has to produce a recommendation.

Which is why the pull only runs one direction. You will likely never get an unrequested line in a draft asking whether this client’s autonomy is being respected. Not because that’s wrong to ask. Because there’s no procedural shape for it to arrive in.

And that asymmetry has a client in it. The recommendation to increase contact treats more care as straightforwardly better. It isn’t, if the person can’t get it in the form, it comes in — and wanting it doesn’t help them if the form is the barrier. What the client will actually accept isn’t a soft consideration sitting next to the clinical one. It’s the condition on anything working at all.

A side note on why it wants to be helpful
These tools are trained to be rated helpful, and people rate agreement as helpful. That’s the short version of what the last three posts worked through, and it’s the machinery underneath everything in this section. If it’s new to you, start there.

The same pressure produces something I haven’t written about before. A response that flags a consideration you didn’t raise rates as more helpful than one that just does the task. So, volunteering the standard concern isn’t a glitch; it’s the behavior these systems were built to have.

Where the additions land

Here’s the part that took me a while to see. Those unrequested additions follow the same density as everything else. Where the territory is common and heavily documented and risk-averse, you get more of them. Where it’s thin, you get fewer.

Researchers have measured the underlying pattern directly: one study tied model accuracy to how many documents in the training data supported a given question, and found accuracy climbed sharply with that count. That was factual recall, not clinical recommendations, so it shows the mechanism working nearby rather than proving my case.

So the tool is most additive exactly where you needed the help least, and quietest where your work is hardest. If your clients sit near the center of all that writing, the additions often land — a consideration you’d moved past, a thing worth a second look. If they sit off-center, the same additions arrive and fewer of them fit.

Two column graphic on AI training data bias the same three unrequested AI additions appear for a client near the center of the training data and a client at the margin They mostly fit the first client and mostly don't fit the second. Banner: same additions, one client they were written for.

The helpfulness and the pull are the same feature, concentrated in the same place. And it’s the implied client again who sits at that center — the one for whom the referral exists, the transportation works, the coverage holds. The clients furthest from him are the ones already absorbing the quiet, concentrated harm in every other part of the system.

What’s written most arrives looking like what’s indicated.

Why you already know this pull

Motivational interviewing named this a long time ago, on the other side of the table. The righting reflex — the clinician’s pull to steer someone toward the correct direction. MI named it precisely so clinicians could catch it in themselves, and harm reduction is the whole discipline of resisting it.

We reach for the more intensive option too. We’re frightened, for the client and for ourselves, and reaching for more is what fear reaches for. Sometimes we jump the gun on it, and sometimes that costs the client something.

So the pull in the output isn’t foreign. It’s the field’s own reflex, arriving from a new direction.

One difference matter, though, and it’s why the skill doesn’t transfer on its own. You catch your own righting reflex by noticing what you want — the tightness right before you correct someone. That has nothing to work with here. There’s no urge to notice, because the pull isn’t yours. What’s there instead is a sentence that reads like the rest of the document, sitting in the middle of your note, with nothing on it to say this is the place your judgment and the standard part ways.

The part that lands in the chart

That line about the full session was the version you can see. There’s a quieter one, in the same note.

When my clinical position sits off the standard, justification comes threaded through the draft without anybody asking for it. Not one defensive paragraph I could see and delete — a clause here, a qualifier there, each small enough to read as thoroughness. Reasons why I chose what I chose. Reasons why I’m not doing what the textbook would say, because real life isn’t that clean.

My standing instructions already say not to do this. It still shows up. I didn’t write any of it, and my name is the one going on it — which is exactly why I read for it.

I didn’t write any of it. My name is the one that goes on it.

To be clear, I’m not saying don’t document. Document the acuity, document what you did about it. What I don’t want in my documentation is a preemptive defense written for a reader who doesn’t exist. If it ever comes to defending the call, the note already does it — that’s what a good note is. Writing the argument in advance costs the narrative, and the client is in there too, being described by a note that’s busy protecting me. That line is mine to hold, and it moves when I’m not watching it.

So, the note argues. My judgment sits in the chart, in my voice, framed as a departure that needs defending — and anyone reading it later meets my reasoning already positioned as the exception. Once the extra is in the record, it doesn’t come back out.

A newer clinician has no read on which is which. They see a draft that sounds more rigorous than what they’d have written alone, and they take the lesson: my clinical reasoning needs defending. It doesn’t. It needs documenting — criteria met, plan, what it addresses, what you’re watching for. Those are different things, and the second one is shorter.

And if you’re pre-licensed, it isn’t only your name on it. Your supervisor co-signs, and they’re reading it as your clinical thinking. You don’t want to be sitting across from them while they ask why you wrote something you didn’t write. The worse version is that nobody asks — it gets signed, it goes in the record, and whatever was in there stays in there. If that’s you, the section below is written for you first.

What to actually do, starting tomorrow

I’m not going to tell you to turn this off. I use it because it saves me real time — writing doesn’t come easily to me and getting what’s in my head onto the page takes me a long while without help. And the unrequested additions aren’t all bad. Call it sixty-forty: more often than not, what it adds is something I’m glad to have, a consideration I hadn’t weighed or a second look at something I moved past too fast.

It’s the other forty that has to be caught, and it arrives looking exactly like the sixty.

But it comes with a cost, and the cost is that everything I get back has two kinds of content in it. What I told it, and what it added. Same voice, same paragraph, nothing marking the seam.

You don’t have to reconstruct what came from where. On a long session you probably can’t, and it isn’t the useful question anyway. The read is simpler: does this match how I want this session documented?

I know how I want my notes written. Strengths-based. Trauma-informed. Culturally responsive. Nothing in it that isn’t clinically necessary, and nothing written in a way that would hurt the client to read beyond what the facts themselves require. Anything that doesn’t meet that comes out, whether it came from me or from the tool.

That list took me years and it’ll keep changing. If you’re early on, you don’t need all of it — you need whatever you’ve got so far, written down where you can see it. Two or three things you know you want every note to do. Your supervisor has a version of this list, and so does your agency; ask what theirs is and start from there. It fills in with time and with the notes you end up wishing you’d written differently.

A side note on models and products
Most of us aren’t using a model directly. We’re using a product built around one, with its own instructions layered on top that neither you nor I can see. Same underlying model, different wrapper, different behavior — which is why your tool may not do what mine does.

Tools differ, and I can’t cover all of them here. So don’t take my experience as a description of yours. What doesn’t differ is that something wrote those sentences and it wasn’t you, so they get read before they get signed.

Who the chart is for

I can’t tell you that what came back isn’t what you wanted. Maybe it is. You hold your own standard for a note and that’s yours to hold, not mine. What I can tell you is that things get added, that the additions carry their own defaults, and that the chart follows the client long after the session does.

Write every note as though the client will read it. Mine have — they request their records for disability claims, for attorneys, for custody matters, or because they just want to know what’s been written about them. But you shouldn’t need it to have happened yet to write that way, and if the thought of a client reading your notes makes you uncomfortable, that’s worth bringing to supervision. When it does happen, every line in there is something you put your name to — including the lines you didn’t write. Not because the tool is dangerous. Because things get in there that could hurt somebody who never chose any of it.

That’s slower than accepting a clean draft, and there’s no version of this where it gets faster. But it shouldn’t cost you time you don’t have. If the tool is doing what it’s supposed to, it already gave you back more time than the read takes, which is the whole reason to use it. So budget the read into your documentation time from the start, the way you’d have budgeted the writing it replaced. It’s the price of the part that helps.


Next: That average isn’t sitting still. What we accept from it is part of what moves it — the larger story I said I’d come back to.

Written with the same tool this post opens with — and argued with the same way.

author avatar
Stephanie Valentin

You can share your post through:

Facebook
Twitter
LinkedIn

Other Posts