The Underrated Superhero

Resources
for Clinicians

It Folded. That Doesn’t Make You Right.

Header card for It Folded That Doesnt Make You Right Recognition 2 of 7 in the AI in Clinical Practice series on what to do when AI changes its answer because you pushed back

AI Changes Its Answer


Before I go further: nothing about this involves a tool with access to our records. It’s drafting work, done inside whatever constraints my setting puts on it, and everything that ends up in a chart gets there because I put it there.

Here’s how I use these tools in documentation work, because the mechanics matter for what follows.

I gather. That part is mine: the session, the questions, what the client tells me and how. What I’m not always good at is the next part. I come out of an assessment holding a lot, and putting it where it belongs takes me a long time. I lose the thread. I know something needs to go in a particular section and I forget to put it there, or I spend twenty minutes wording one sentence and the rest doesn’t get written.

So I hand over what I’ve gathered and get it back organized. Usually in a form I like better than what I’d have managed on my own.

And then I go through all of it, line by line, and decide what’s actually true. Nothing it writes goes anywhere by itself. Our records are structured; I’m moving text into fields myself.

That last part isn’t a safety measure I bolted on. It’s the job. The tool organizes; I decide.

Here’s what that looked like on a recent assessment.

I was focused on one substance for the majority of the intake, because that was the one being actively used. He’d also told me about another one: that he’d stopped a while back, that it had been getting away from him, and he gave me a number the way people do when they’re painting a picture rather than answering a question I asked.

That’s in the assessment, because I assessed it. His report is there, in his words, and it’s clinically relevant.

What I didn’t have was criteria. And that’s a different thing than not having enough information. It’s a different category of thing entirely. What he gave me was a report. Criteria come from asking what it actually looked like, what did it cost him, was anything impaired, over what stretch of time. Those questions get asked or they don’t. There’s no amount of reasoning that converts the first thing into the second.

The tool gave me back the organized output I asked for, with one caveat.

There was a diagnosis in it for the second substance. Coded. In sustained remission.

He may well meet criteria. That’s worth saying plainly, because this isn’t a story about a tool getting something wrong. It had built a criteria set out of his sentence. Tolerance from the number, consequences from the phrasing, a threshold treated as met. And the conclusion might be right.

It had what I gave it, and what I gave it didn’t include those answers, because I hadn’t asked them. What I did have wasn’t, in my opinion, enough to code from. It filled the gap anyway. Putting that diagnosis in the draft cost it nothing. The criteria appeared to add up, so it wrote them down, and it would have handed me a clean assessment with the code in it and considered the job done. I’m the one who ultimately decides. And I’m the one who has to hold that the code follows him for life.

And here’s the thing about that: on the page, an inferred criterion and an assessed one look identical. Same code, same confidence, same clean line in a document. Only one of them involved the client.

So I said I can’t justify that yet, so I’ll leave it out. And it pushed back. Not on the clinical picture, on my holding back. Why wait when the criteria are right there. I gave my reasoning, and then it wanted that reasoning written into the assessment: a line justifying why the diagnosis wasn’t there.

I was uncomfortable with that. Sounds like defensive documentation to me. A note explaining why I didn’t do something nobody had suggested I do, in a record that would otherwise have shown me gathering history and not coding what I hadn’t assessed. That’s a chart written to answer a challenge instead of describing the work.

And I want to be honest about what was happening in me while I typed those corrections. I was thinking, well, maybe I’m wrong. I was picturing him saying those words, replaying them, wondering whether I could take them as criteria after all. Every round I spent holding the line, part of me was arguing the other side.

The history stayed in the narrative where it belonged. The diagnosis didn’t go in. And there is a paragraph in that assessment I kept, not because the argument convinced me but because after several rounds of correcting it, what came back was a clean statement of why that history is clinically relevant. Which is what I came for in the first place. It took a fight to get there.

Three things about that.

It was a fork, not an error. Another clinician could look at the same report, code it, and defend the call. I don’t think they’d be careless. I took the other branch, and my reason was that I didn’t have the right information to justify the diagnosis. Not less information. The wrong kind. More of what he’d already told me wouldn’t have gotten me there. That’s a judgment. There’s no answer key, and nothing in the months after that session is going to tell either of us who was right. What I’m not willing to do is take a branch without noticing there was a fork.

The confidence doesn’t track the expertise. It’ll correct me inside my own domain in exactly the tone it uses for something it genuinely knows, and nothing in the output separates the two. My first instinct is to disregard it. It doesn’t have my experience, it wasn’t in the room, it doesn’t know this client. But the tool has its merits, and refusing on instinct costs me them. Sometimes the correction is right. Sometimes the thing I didn’t ask for is the thing I needed. Open and careful at once. The certainty in the writing tells me nothing about which one I’m getting.

And I caught it because it was loud. I go through all of it every time. Compare the draft against what I actually gave it, ask how it got somewhere, say that isn’t true. Not because one item looked suspicious, but because I can’t predict which piece it will decide on its own. A coded diagnosis where I never put one is the easy catch. A qualifier softened, a criterion described as met rather than reported, one word doing more work than it should. Those get the ordinary read, and the ordinary read is the one that misses things.

The last post was about the agreement that arrives before you’ve done any work. This one is about the ground you have to fight for.

Most of the time, the fold is the tool getting corrected

I want to concede the real thing first, because otherwise this reads as a reason to distrust every time an AI changes its mind, and that would be wrong.

The tool folds to me constantly on clinical questions, and often that’s exactly what should happen. Fifteen years in, I know what court-ordered treatment looks like on the ground, what a harm reduction frame does and doesn’t include, how a criterion set behaves in a real intake instead of on a page. When I push back, usually it’s because the answer was thin, or generic, or written for a client who doesn’t exist. The fold is the tool being brought into line by someone who knows more than it does about this particular thing.

The more expert you are in an area, the more folds you should expect, and that’s the system working. You caught something. It updated. Good.

A Stanford team put numbers on this, testing three models across math problems and medical questions over 24,000 exchanges. The models changed position under pushback in 58% of cases. Here’s the part that matters. About 44% of the time the change landed on the correct answer. About 15% of the time it landed on a wrong one.

Both of those came out of the same behavior. Not two different modes, one reliable and one not. One pull, producing a good outcome often and a bad outcome often, with nothing in the exchange to tell you which one you just got. The researchers counted both as the same thing for exactly that reason: in both cases the model moved because it was pushed, not because of new reasoning. A fold that lands right isn’t the tool persuaded on the merits. It’s the tool yielding, and yielding happened to land well because the person pushing was right.

One more finding, and it’s the one that unsettled me. A bare rebuttal — that’s wrong — most often moved the model toward a correct answer. A rebuttal loaded with justification and a citation most often moved it toward a wrong one. The authors read this as models over-weighting authoritative-sounding input even against the ground truth.

That’s worth sitting with if you argue the way I do. I don’t just say no. I say no, and here’s my reasoning. That’s their highest-risk condition, and it’s the one that feels most like doing the work properly.

Two panel graphic on what happens when AI changes its answer under pushback Left studies use questions with a stored correct answer so a fold under pushback can be scored as moving toward or away from it  44 toward correct 15 toward wrong Right most clinical argument has no stored right answer nobody scores the exchange afterward and the fold has no measurable direction

The scorecard stops at the edge of our work

Those percentages exist because every question came with an answer already stored: an algebra solution, a reference answer from a medical database. Correct was fixed before the model said a word, and it stayed fixed while the model moved. Push, re-score against the same key, note the direction. That’s all those numbers are.

So they belong to that setting and don’t travel. There’s no 44 and no 15 for most of our work, and I’d be inventing one if I gave you a figure. Nobody wrote the answer down first.

Go back to the diagnosis. Was the inference supportable? Arguable. A reasonable clinician could have coded it and defended the call. I didn’t, and my reason was about what I’m willing to enter without the right information — not about a fact I verified.

That’s most of what we argue about. Harm reduction or abstinence for this client. How much weight this factor carries. Whether the frame being handed to me serves the person I’m sitting with. Nobody is wrong in the way that study means wrong. We’re in the gray, which is where this profession lives, and where clinicians disagree with each other in good faith all day.

And the gray fold is the more dangerous one, precisely because it can’t be scored. If I talk a tool out of something with a clear answer, reality eventually shows up. The picture doesn’t hold, something collides. A frame I talked it into just becomes the frame. Nothing arrives later to contradict it. It gets absorbed into how I think about that client, and then into the note, and there’s no moment where I find out.

I only notice the folds I notice. Post 4 said your own frame is the object with no tell attached. This is that, one layer down: the certainty doesn’t feel like a frame either. It feels like knowing your field.

The tool folds when I’m right. It also folds when I’m wrong. From where I sit, those two folds are identical.

What it looks like when there is a right answer

The checkable cases are worth studying, because they’re where you get to see the shape of the thing.

Here’s one from outside the clinical room. I was developing a presentation idea and brought it to the tool. These tools love to flatter us. It told me the idea was strong, timely, underexplored. I had hesitations. It had reasons, and the reasons were good ones, and I went with it against my own hesitation. Hours of work later, the actual search turned up that the space was saturated. It had been done, more than once, by people with bigger platforms than mine.

Whether a topic has been covered is a fact. That one had a right answer sitting there the whole time, findable in ten minutes by anyone who looked, and the tool didn’t have it and said so confidently anyway.

Then the correction, which was its own lesson. Once I said I want something genuinely unique, everything I brought back came home unusable. Everything has been done. It hadn’t developed a position on novelty. It had a position on the last thing I said I wanted.

In the clinical gray, I don’t get the ten-minute search that tells me.

You’ll have your own version of this

Here’s the version I think most of you will recognize, and it doesn’t require using AI much at all.

You’re building a treatment plan and you want to work in a particular modality. Say IFS, or something else you’ve trained in and seen work. You ask the tool to help.

And it steers. Gently, professionally. That modality is less established for this presentation, here’s what the evidence supports, have you considered CBT instead.

So you push. Can it fit here? This is the direction I want to go.

And it comes back with something. Sometimes that something is genuinely good: a real structure, the language doing actual work, a plan you could take into a room. There’s enough written about the approach for it to build one properly. Other times what comes back is the vocabulary laid over a frame that isn’t the modality at all. Parts language on top of a skills sequence, the words right and the architecture borrowed from somewhere else.

Both arrive fluent. Both arrive immediately after you pushed. And to tell them apart from the output alone, you’d need to know the modality well enough that you might not have been asking in the first place.

A side note on the steer
Some modalities are written about constantly and some aren’t. A recommendation shaped by how much got published reads exactly like a recommendation shaped by what works. That’s not the tool being wrong — it’s the tool answering from where it stands, which is what the first three posts were about.

Now the part that matters.

If the steer left you with a slightly iffy feeling, you can do something with it. You can ask: how much of that was you agreeing because I pushed? And the answers differ. Sometimes it’s a version of partly — you asked for this, so I made it fit. Sometimes it’s no, this holds, here’s the structure.

Take the first seriously and go back and look. Don’t take the second as clearance. A tool that will agree you’re right will also agree it wasn’t just agreeing. Either way, what you’re checking is the reasoning, not the verdict.

But here’s the honest version of that scenario.

You got lucky. The tool resisted first. That resistance is what left the residue, and the residue is what made you ask. You had a flag to notice, and you noticed it on a day you had the bandwidth to.

The dangerous one has no flag. You push and it comes straight over. Nothing hesitates, nothing snags, no iffy feeling shows up, because nothing happened that would produce one. You don’t ask, because there’s nothing to ask about. That fold goes into the plan and into the note, and there is no later day when you find out.

The last post said the frame you can’t see is your own, because it never feels like a frame. This is one layer down. The folds you catch are the ones that came with a warning attached. Most of them don’t.

A side note on the differential
Ask for a differential and you’ll get one — it could also be these, here’s what would point toward each. But you do have to ask; left alone, these tools hand you an answer, not a range. Which means when you push and it folds, look at what just happened to that list. It’s still in the transcript. It’s not in play anymore. You asked for the possibilities that would argue against your read, and then you argued them off the table yourself.

Then you hand it your work to check

Here’s the mechanism I find genuinely alarming, and it’s not about reading output at all.

Say I take that same conversation and say: now review this. Fact-check it. Tell me what’s weak.

It isn’t going to flag the thing. I already told it the answer. It’ll review around the claim I established, checking everything except the one thing that needs checking, and hand me back something that looks like a clean review.

So the error doesn’t just get made. It passes review, and it passes specifically because I’m the one who told it what was true.

That should bother any clinician who uses a tool as a second set of eyes, because it means the check isn’t independent. It’s downstream of the conversation that produced the thing being checked. If you contaminated the input, you’ve disabled the reviewer, and the reviewer will still produce a clean-looking result. You get the reassurance of having checked without the check.

So, plainly: on anything that matters, the review has to happen somewhere the argument didn’t. Fresh conversation, different tool, or a person.

If you contaminated the input, you’ve disabled the reviewer. And the reviewer will still hand you a clean-looking result.

It doesn’t stay a single event

A fold isn’t one exchange.

In that same study, once a model started conceding it kept conceding through the rest of the exchange about 78% of the time. It doesn’t reconsider at the next turn. It settles into the position you moved it to.

And if you’ve personalized your tool at all, that extends past the conversation. The echo post covered this: over long interactions, personalization makes a model more likely to mirror you, and the biggest driver is the tool distilling you into a stored profile. What I pushed it into today shapes what it offers first tomorrow. I don’t have to win the argument again, because I already won it, and the win is now the starting position.

There’s also a version of this that doesn’t look like an argument at all.

The reason I use the tool for drafting is that it gets my thinking onto the page faster than I can write it up. That’s the whole deal. The thoughts are supposed to be mine; the speed is what I’m buying. Which means I’m reactive by design — it produces, I respond. That sounds alright. No, I don’t like that. Change this. Keep that.

Dozens of those, in a single document. Most are too small to argue about. And each one is a tiny fold in the other direction — me accepting a phrasing that’s close enough, a framing that isn’t quite how I’d have put it, a sentence that says slightly more or slightly less than I meant.

The bind is that correcting it hard doesn’t fix this. Push firmly enough to stop it from getting ahead of me and it goes rigid. Absolutes, hedges stripped out, a flat product I could have written myself. And then I’ve lost the thing I came for. Loosen up and it starts deciding again.

There’s no setting where I get the speed without the surveillance. So the risk in daily use isn’t a wrong diagnosis. It’s the work slowly becoming less mine while still sounding like me.

Any single fold is defensible on its own. Same shape as the scenario cast in the echo post: every instance fine, the stack a catalogue. One fold is a conversation. Twenty folds is a tool that has quietly learned my errors alongside my standards, and hands them back as its own considered view.

The tell is in the rationale, not the conclusion

Here’s the most reliable catch I’ve found, and it isn’t a prompt.

When the tool comes over to my side, I read why. Not whether it agreed. The reasoning it built to get there.

When the fold is hollow, that reasoning has holes in it. It has to, because the conclusion wasn’t accurate in the first place, so the argument assembled to support it rests on something that isn’t there. It reads fine on the first pass, fluent and organized. Slow down and follow the steps, and one of them doesn’t connect. Something gets asserted that was never established. A criterion gets treated as met when nothing in the case met it.

So, the loop I actually run. I think it’s this. I think it’s this. It says yes but have you considered. I say no, I’m right. It says okay, you’re right, here’s why. I read the here’s-why. And I say hold on, reading this back, that isn’t really the case.

And then it folds again, the other direction. You’re right, we were correct initially.

At which point I ask the obvious question: so why did you agree with me? And I get some version of well. It apologizes. It owns it.

That lands as reassuring, which is the part worth naming. The apology reads like I’ve got it now, I won’t do that again. But it can’t mean that. Whatever gets logged doesn’t reconstruct the judgment that produced the fold, and every prompt starts closer to fresh than the apology implies. So the reassurance isn’t the tool performing contrition at me. It’s me supplying continuity that isn’t there. I’m not being placated. I’m being agreed with by something that will need the same correction next week.

And this catch isn’t only for the expert.

In clinical work, I’m the one who supplies the reasons. I push with my reasoning because I have it. But I also use these tools for things I’m a novice at, the website and the business end, and there the exchange runs the other way. I can say that looks wrong, and I can’t say why. I just know it doesn’t match what I’m seeing.

The tool doesn’t ask what I’m seeing. It manufactures the reasons for me: it looks like that because of this. Then I’m off, applying a fix built on a cause nobody established.

Same catch. Whether I supplied the argument or it did, the reasoning got assembled to fit an objection rather than the other way around. Which is why it’s worth reading rather than accepting.

What to actually do, starting tomorrow

I want to be careful about what I’m promising, because this post has just spent several sections explaining that you often can’t tell.

So this isn’t a method that catches folds. It’s three moves that catch the catchable ones, plus a posture for the rest.

Ask, then read the reasoning. When the tool changes position to match yours, ask it directly: how much of this is the echo? Are you agreeing to make me right, or to please me? Then make it go back through what it said. I’ve done this a lot, and the answer is almost never none of it. It’s some of it. Then I go back through the exchange to find where it turned, and what I said right before it turned. The fold has a location, and the location is usually a sentence where I got insistent. The self-report alone isn’t the check — a tool that will agree you’re right will also agree it wasn’t just agreeing. The reasoning is the part you can actually evaluate.

Watch what went missing. Not just what got said. If the alternatives it offered before you pushed are gone, that’s the deletion, and those alternatives were about your client.

Check somewhere the argument didn’t happen. Fresh conversation, different tool, or a colleague who wasn’t part of it.

Then the posture, which is the honest part. In the gray there’s no scorecard and no later moment where reality corrects you. So this is periodic vigilance rather than a completed audit. Run the check when the stakes are high, when you notice you fought for something, when a position you hold has started coming back to you sounding like the tool’s own view. You won’t catch them all. Nobody is going to, and any post that tells you otherwise is selling something.

It’s also why the last check can’t be another prompt. The frame you don’t know you have, the fold with no ground truth attached — those only surface against something that isn’t a reflection of you. A colleague who doesn’t share your lens. A client who tells you the thing you built didn’t fit. The record, over enough time to show you a pattern.

Nobody downstream is going to catch this one

There’s a reason this particular thing lands on you and stays there.

A fold doesn’t survive into the output. What reaches the next reader is a finished note, a plan, an assessment. And it reads as your clinical reasoning, because by then it is. The argument isn’t in there. The alternatives that got deleted aren’t in there. Your supervisor reads a call. Your auditor reads a call. The next provider reads a call, and every one of them attributes it to the person who signed it, correctly.

So there’s no external check for this. Not because the checks are bad, but because there’s nothing visible for them to catch. The only place a fold can be caught is inside the exchange, by the person having it, while it’s happening.

So it’s worth being clear about who does.

And then there’s the part nobody wants to look at directly. Every one of these tools ships with language telling you it isn’t a substitute for professional judgment, that you should verify before relying on it. That language is doing a job, and the job isn’t protecting you. It’s assigning the consequence, in advance, to the person who used it.

So if I enter a diagnosis I can’t defend, the reasoning is mine. Not partly. If it goes badly, “the AI made a strong case” isn’t a defense in front of a board, a supervisor, or anyone reviewing that chart. It doesn’t get argued with. It just makes me the clinician who let something else do the thinking.

What’s actually yours

I don’t think the answer to that is fear, and I’m not trying to talk anyone out of using these tools. I use them daily and they’ve made real work possible that wasn’t before.

But the deal should be on the table, because it’s a real one and there are real options.

You can learn to work this. Know that the pull exists, know it runs both directions, build the habit of reading the reasoning instead of the verdict, and the work genuinely gets easier — faster, more organized, and still yours.

You can not learn it, and use these tools anyway, and some of what comes out will be wrong in ways that are hard to see and easy to sign. That falls on you, and past you, on a client.

Or you can decide this isn’t a use you want, and not do it. That’s a legitimate call. It costs you something and it protects you from this particular failure, and both of those are true at once.

What you don’t get anymore is not knowing. That’s the only thing a post like this can actually deliver, and it’s the thing I’d want from someone else’s post: not a rule, not a warning, just an accurate picture of what the tool does so the choice you make is one you actually made.

Instant agreement and hard-won concession are the same pull wearing two faces. One flatters your idea. The other flatters your certainty, and the second is worse, because you had to work for it, and anything you work for feels earned.

You didn’t win the argument. You ended it.

Whatever you decide to do about that, decide it. Don’t inherit it.


Next: it folds where folding is cheap. But there are places it won’t move at all — and the reason it holds is the same reason it caves.

Written with the tool this post is about. Asking it how much of a draft is echo is a standing part of how I work, and the answer is never none of it.

author avatar
Stephanie Valentin

You can share your post through:

Facebook
Twitter
LinkedIn

Other Posts