AI in clinical documentation
There is a question I hear about AI in clinical work more than any other, and it goes something like this. It keeps getting better, doesn’t it? At some point I won’t have to think this hard about it.
I want to take that seriously, because the observation underneath it is correct. AI capability is improving. Genuinely. The output from a year ago and the output today are not the same document, and a year from now it will be better than today. Anyone telling you otherwise hasn’t been paying attention, and I’m not going to be that person. So let me say it plainly. Yes. It is getting better. It will keep getting better.
Before I go further, one thing to put on the table. This post is for clinicians using AI in their work, and that covers a wide range of people for a wide range of reasons. Some of you chose it deliberately. Some of you had it handed to you when your agency rolled something out. Some of you reach for it because you are newer and trying to feel competent — which is a feeling a lot of new clinicians carry for reasons that go well past AI. Some of you reach for it because you are senior and trying to stay sane.
The reasons are not the same, and they don’t all carry the same weight, and this post is not going to sort them. The post is about what happens to the output once it is in front of you, regardless of how it got onto your screen.
The entire post turns on a small word inside that question. Better. Better at what.

Two Axes, Not One
When clinicians say AI is getting better, they almost always mean one specific thing. The output is more fluent. More accurate on the verifiable parts. Less obviously clumsy. Fewer of the embarrassing mistakes that a year ago would have made you screenshot it for a friend. All of that is real, and I’m not going to argue any of it.
But notice what that improvement is measuring. It is measuring capability — how well the AI does the surface job. How well it writes a sentence. How well it structures a paragraph. How well it sounds like a thoughtful colleague. Capability is one axis, and it is improving.
Capability is also not the only axis. There is another one, and the work this blog is about lives on that one. The second axis is standpoint — what the output assumes, who it imagines as the default and who it doesn’t imagine at all, what frame the language quietly carries from whose data went in. Every AI system has a standpoint. It can’t not have one.
A system that learned from no human data would know nothing. The only systems that exist are ones that learned from someone, and what someone wrote down is always a position. There is no view-from-nowhere AI sitting just over the horizon, and there isn’t going to be, because that is not how learning from data works. Standpoint is structural to what AI is.
Capability and standpoint are two different axes. Capability is improving. Standpoint is not disappearing, because there is no zero-standpoint version of anything that learned from anyone. And here is the part that catches people off guard, in my experience. A more capable AI is not a less standpointed one. It is a more fluent one. Sometimes a more persuasive one. Often a smoother one. The kind of output that two years ago would have been clumsy enough to catch is now polished enough that you have to slow down to find what’s bent in it.

Better, in the way most clinicians mean it, can quietly mean harder to see.

A Pattern In The Work
There is a clinical version of this that comes up often enough to be worth naming.
A client comes to treatment fluent. A second round of treatment, or a third. They know the language. They can articulate their own patterns before you offer them, they can name their attachment style, they describe their cravings in shapes you could chart. Compared to a client in their first week of treatment — confused, defended, blunt — this client is, on every measurable surface, more capable.
What tends to come next is the part of the pattern worth attention.
The fluent client is often harder to read, not easier. The articulate language can be doing the work of not feeling the thing, or of giving back the version of the answer the last clinician seemed to want. The capability went up. The clinical work did not go away. It moved. It got harder. And a clinician under pressure can absolutely mistake the new fluency for the work being done. That isn’t a failure of judgment. It is a failure of the conditions the clinician is working in.
That is exactly the shape of the AI problem. Polish on the capability axis is not progress on the standpoint axis. The two axes do not touch.

What this looks like in real life
A few therapist friends have told me some version of this lately. The cleanest example: one of them was excited about transcription software her agency had rolled out. It was helping her document faster, and documentation had been eating her evenings. I was happy for her. That is a real win, and the time it gives her back is time she can spend doing the actual clinical work, or, you know, sleeping.
I asked her, in the way clinicians ask each other things, whether the output was trauma-informed. Whether it read as strength-based. Whether it was culturally responsive.
Some version of: I don’t know. And honestly, I don’t have time to check for that.
I want to be clear, because this part matters more than anything else I am going to say in this post. None of these friends are doing anything wrong. They are careful clinicians working under conditions they did not design, with workloads they did not choose, in systems that hand them tools and don’t also hand them the time to use those tools well. They are surviving the week. I have been right where they are. The post is not saying they should be doing more.
The post is saying something narrower. The polish of the output is not evidence that the checking has been done. A tool that gets faster and smoother and more pleasant to use is doing well on the capability axis, and that improvement is real. It just is not, by itself, evidence of anything on the standpoint axis. Which is the axis the questions in the previous paragraph actually live on.
Here is the part the friends in my story haven’t slowed down to see, because they haven’t had time to. A lot of the documentation AI on the market right now produces notes that come back fluent, clinically literate, easy to sign. Read one closely and you can feel what it is actually oriented toward. Getting paid. Surviving audit. Satisfying medical necessity. Compliance is doing a lot of the work in the standpoint.
Trauma-informed, strength-based, culturally responsive — those are the questions a clinician might have wanted the tool to be careful about, and they are not what the tool was built to be careful about. The defaults reflect what the training data was for, and a lot of training data is charts that had to clear insurance review. Some vendors are working on this. Some aren’t. None of them have solved it.
So the output ends up fluent at the thing it was built for, and quieter on the things you would have wanted it to be careful about. That isn’t the software malfunctioning. That is the standpoint axis, doing exactly what it does, in language polished enough that nobody had a reason to slow down on it.
There are bigger fights inside what I just said. Whether agencies should be deploying tools whose training optimizes for compliance over care. Whether the insurance review structures behind those defaults are themselves the standpoint problem upstream of the AI. Whether the time pressure that makes I don’t have time to check a rational answer is itself the thing that needs addressing. Those are real. This post does not take them on. It takes on one layer underneath — what the output is actually oriented toward, and what becomes possible once you can see that.

What you’re entitled to know
I am not telling you what to do with any of this. That is not the job of this post. The decision about what you read carefully and what you let go belongs to you — to whoever is doing the work in your particular setting, with your particular workload, on a particular Tuesday. Nobody reading over your shoulder gets to make that call for you.
What I think you are entitled to, before you make that decision, is to know what the decision is actually about. The polish you are seeing is real, and it is on one axis. The questions you would have wanted to ask — trauma-informed, strength-based, culturally responsive, accountable to the people you actually serve — are on the other. Whatever you decide on a hard Tuesday, it is not the same decision when you know how the axes work as it was before you knew. That is what this post is for. Not to make the choice for you. To name what the choice is.

What comes next
So far this post has been about the structural reason capability does not retire your judgment. There is a harder layer underneath, and it is the one I sat with longest before writing any of this.
I have been telling you that AI output needs careful reading because the standpoint axis does not improve no matter how good capability gets. There is a version of the problem that is harder than that. Even when a clinician does read carefully, even when the clinician is paying full attention, the reading itself can be quietly steered. The AI bends toward whoever is talking to it. Output that confirms what the clinician already believes, that sounds like the position the clinician would have taken anyway, that reads as good clinical thinking because it matches — that output doesn’t get caught. Not because nobody is looking. Because nothing in it is asking to be caught.
The catch has to fire on agreement, not just on friction. That is the next post. It is the one that matters most.

Next: the AI output hardest to catch is the one that feels right.