The Underrated Superhero

Resources
for Clinicians

It’s Not Wrong. That’s the Problem.

AI Sycophancy


I asked an AI to build me a set of clinical scenarios with a diverse cast, and I was specific, because I know what gets left out. Vary the race, the gender, the body, the family structure, the class. It did, on the surface. But a few scenarios in, I started seeing the seams, and they weren’t one-offs. The Black fathers were incarcerated or absent, not once but across scenario after scenario. The Black girls kept turning up as the larger-bodied ones. The Latino characters came pre-loaded with immigration stories I hadn’t asked for, again and again. The Asian women wore glasses. Each instance was defensible alone. Stacked up, they were a catalogue: diversity on the surface, the same worn stereotypes underneath every time a group actually showed up.

And then there was who didn’t show up. No queer characters unless I named them, and even when I did, they’d land as the protagonist and disappear from the supporting cast, as if the background of the world defaulted straight. The stereotypes I could at least see once I looked. The absences I had to go hunting for.

That output was more capable, not less. A year ago the AI would have handed me a cast of quietly white-coded characters and called it diverse. This version actually produced Black characters with specific lives. The capability went up. And the added capability is exactly what delivered the stereotype. When I asked for a Black father, it reached, over and over, for the same story. I watched it default to one narrative when countless were available, and do it fluently enough that the narrowing looked like richness. The richer detail didn’t make the problem easier to catch. It made it harder, because a stereotype arriving as vivid, specific detail looks like good writing.

I caught it, but it took a while. I’d done what you’re told to do: prompt carefully, name every dimension, leave nothing to chance. Vary race, gender, body, family, class. And it checked every box. The cast varied on everything I’d asked for, and I was impressed, because the prompting had clearly worked. That was the problem. Being impressed that it did what I asked is exactly what kept me from looking at what it actually gave me. The real subject here isn’t bias. It’s that the output agreed with my request, and the agreement is what stopped me looking.

The real subject here isn’t bias. It’s that the output agreed with my request, and the agreement is what stopped me looking.

The dangerous output is the one that feels right

Most of what gets said about catching AI mistakes is about friction. You read the output, something rubs you the wrong way, you stop and check. A diagnosis that’s too aggressive. A phrasing that isn’t trauma-informed. A scenario that reaches for an obvious stereotype. The output rubs against what you know, and the rub is the signal.

That catch is real, and it matters, and it’s also the easy one. Everybody already knows to double-check the thing that looks off. You don’t need a practice for that. You need a practice for the opposite case: the output that doesn’t rub against anything, because it matches what you already thought.

When the AI hands you something that lines up with your own clinical instinct, you don’t experience that as the AI being shaped by your input. You experience it as the AI being right. The agreement reads as confirmation, an outside voice, fluent and confident, telling you what you already believed. And the moment it agrees, your scrutiny quietly switches off. Why would you audit the thing you already agree with?

That’s the echo. And notice it isn’t about whether the output is right or wrong. In the scenario case the characters were real, specific, well-written, technically fine line by line. The failure wasn’t a false statement I believed. It was that the parts matching what I’d asked for sailed through unexamined, because they matched. Right or wrong, I’d stopped evaluating the moment it looked like what I wanted. Agreement is what switches off the check. Correctness has nothing to do with it.

Agreement is what switches off the check. Correctness has nothing to do with it.

And this is the turn the title is pointing at. None of those outputs were errors on the AI’s terms. The stereotyped father is faithful to the story its training data tells most often. The default-straight supporting cast is faithful to a world where straight goes unmarked. Each one is right, relative to the standpoint it was built from. They’re only wrong for the client in front of me. The mistake isn’t that the AI malfunctioned. It’s that I’m not obligated to hold its standpoint just because it handed me the output fluently, and the fluency is exactly what tempts me to.

The uncomfortable baseline underneath all of it: the AI tends to hand back what you want to hear. The field has a name for this, sycophancy, and a reason for it: the model was trained to be rated helpful, and people tend to rate agreement as helpful. Developers actively work to reduce it, and newer models are better than older ones, so the degree is not fixed. But it does not drop to zero, for the same reason capability and standpoint were two different axes in the last post: a model that never bent toward the person it was talking to would be a different kind of thing. The pull gets smaller. It does not leave.

I call it the echo, because the word that helps me hold it isn’t about the model’s behavior in the abstract, it’s about what that behavior does to the clinician on the other end: your own frame, handed back fluently, mistaken for confirmation. Not maliciously, not always, but as the pull, the path of least resistance. Ask it to justify the call you’ve already made and it will, fluently. That’s the default the whole post is about.

A note before I go further, because the next stretch is going to sound like it’s only for heavy users. Some of you barely touch AI. A reworded email here, a colleague’s tool there, a note draft you skim and sign. You might already be deciding this isn’t your problem. Stay with me. Most of what follows is about the echo you build through heavy use, because that’s the version I know best and the one I can show you from the inside. But there’s a quieter version that reaches the once-in-a-while user too, and I’ll get to it. It turns out the less you use the tool, the less reason you have to suspect it has a slant at all. The heavy user’s risk is loud. Yours is silent. Neither of you is exempt.

And if you’re a light user, hold onto one thing even if the rest of this reads as someone else’s problem. The echo is what a standpoint does to you over time. But the standpoint comes first, and it’s there from your very first question: the tool has a position, and what you ask shapes what it hands back. That’s the foundation, it’s where this series started, and it’s the piece to read if this one isn’t quite yours. The echo is the advanced version of a problem you already have the first time you open the thing.

This isn’t an argument against personalizing your AI

The obvious read of all that is “so keep the AI at arm’s length, don’t let it learn your preferences, stay suspicious.” That’s not what I’m saying, and I think it’s wrong.

I personalize my AI heavily, on purpose. Over a lot of sessions I’ve taught it the standards I hold: that I diagnose conservatively, that I want strengths-based language, that I don’t take abstinence-only framing as the default, that I read a note for whether it would shame the client if they saw it. Baking those in is good. I’m not re-teaching the same things every session. The working relationship gets sharper, and I trust it. I can train the AI, to a real degree, to be my eyes, to carry my lens into work I don’t have time to do from scratch.

But here’s the catch in being trained to be my eyes. It never becomes my eyes. It has its own standpoint, so it can’t see everything I can see, and it can’t stop seeing from where it sits. Training moves its frame closer to mine. It doesn’t give it mine. So two gaps run at once. It can’t see what I see: the years, the client’s face, the thing in the room that never made it into the prompt. And it sees from somewhere else, an aggregate built from its training data, sitting near the statistical center, which means it’s weakest exactly where my clients are furthest from that center.

Put those together and the benefit and the limit turn out to be the same fact. I can extend my sight through the tool, and what it extends is partial in two directions at once: short of what I see, and tilted toward where it stands. The training narrows the gap. It can’t close it, because closing it would mean the AI having no standpoint, and there is no such thing.

Which is why the structural fact about personalizing is almost the opposite of “personalize less”:
The better it reflects you, the more its agreement sounds like your own clinical reasoning, so the harder it is to feel the difference between a second opinion and your own voice played back. That’s the echo, and personalizing makes it louder, not softer.

The echo doesn’t quiet down as the relationship improves. It gets more seamless, and the places where the AI is just reflecting me back get harder to tell from the places where it’s actually adding something. Diligence doesn’t trade off against personalization. It scales with it.

What it looks like when being my eyes works against me

Abstract enough. Here’s what the echo actually looks like in the work, getting harder to catch as it goes.

The leaked residue. I’ve trained the AI toward harm reduction. It’s central to how I practice. Ask it for harm-reduction content on substance use and it delivers. But the aggregate underneath it sits closer to abstinence, because that’s what dominates the writing on substance use. So the frame comes back right and the residue comes back wrong: a stray “clean,” a relapse framed as failure, a quiet assumption that reduction is a way station to abstinence rather than a legitimate place to stand. I asked for harm reduction. I got harm reduction with stigma in the connective tissue. And because the piece is mostly what I wanted, the agreement-signal fires and the leaked word rides through, unless I’m reading for it.

The collapse where the data runs out. Now take harm reduction out of substance use and into a realm where it’s been pushed to the margins: harm reduction around self-injury, around eating, around risk behavior the dominant discourse only ever discusses in terms of stopping. The work exists. There are clinicians and researchers doing it, and doing it well. But it’s under-attended, under-funded, talked about and favored less, so it’s thin in the AI’s training data. Not because the approach is unproven, but because the data inherits the same neglect the field has lived with. Ask for a harm-reduction frame there and the AI has almost nothing to draw on, so it falls hard to the dominant frame of control, elimination, abstinence, and hands it back wearing harm-reduction vocabulary, sounding just as fluent as it did on solid ground.

This is the part worth sitting with. The AI is least reliable exactly where the discourse has been most dismissive, which is often exactly where the most important boundary-pushing work is happening, and where I most need it to think differently. The places I most need a second mind are the places it has the least to think with.

The silent omission. The hardest one, because there’s nothing on the page to catch. I’m trained toward trauma-informed care. Hand the AI a justice-involved client and don’t name carceral trauma, and it won’t volunteer it. The note comes back competent, trauma-informed in the ways I flagged, and silent on the exposure that matters most for this person. Nothing rubs. There’s no wrong sentence to find. The miss is an absence, and absences don’t trip the agreement-check, because there’s nothing there to agree with. If I’m not bringing the lens, the AI won’t bring it either, and the client who absorbs that silence is exactly the client I built my work around.

Three catches, getting quieter: a stereotype you can see if you look, a residue you can feel if you read for it, an absence you can only catch if you already knew to ask.

Better prompting raises the ceiling. It doesn’t remove it.

Every clinician reading this is reaching for the same fix I reached for. If it bends toward me, I’ll just prompt it to push back. Tell it to argue the other side. Tell it to flag where it’s defaulting. Tell it what to watch for.
Do it. It genuinely helps, and it’s a real skill. “Give me a harm-reduction frame and flag where you’re sliding to abstinence because the literature is thin” surfaces more than not asking. Wording the request open instead of closed (I’m thinking MDD, but I’m genuinely unsure; what would push toward adjustment instead rather than justify MDD for this client) keeps a door open that “justify” slams shut. That’s controllable, and it matters, and you should do it every time.

But it has a ceiling, and the ceiling is the whole point. When I ask the AI to argue against me, the argument it generates is its model of the counter-argument, built from the same aggregate, tilted the same way it was already tilted. In a well-documented area that proxy is decent. In the data-poor area, the “pushback” it produces is hollow, because it’s working from thin, lower-quality material for a position it has barely seen. So the prompt that works beautifully on a well-charted question gives me a confident, empty counter-frame on the exact question where I needed a real one most. Better prompting raises the ceiling, and raises it least where my data is thinnest, which is where my actual work lives.

There’s a tell in this worth learning to read. Notice which prompts you have to fight for. Some prompting is just opening a door your own framing closed (I’m thinking MDD, what would push toward adjustment), and that’s about your frame, not the literature. Both calls are well-charted; you’re only countering your own lean. But when you’re pushing hard to get the AI to hold a frame at all, when it keeps sliding back no matter how you word it, that resistance is itself a signal. It usually means the frame is thin in the training data: under-researched, contested, pushed to the margins. The effort you’re spending tells you how far off-center you’ve wandered. And off-center is exactly where the AI is least reliable and your own judgment carries the most weight. The harder you have to work the prompt, the less you should trust what it finally gives you.

And the deepest version: a prompt can only open a door I know is there. I can ask “what am I missing culturally?” but only if it’s already occurred to me to ask. The miss that hurts the most is the dimension I never thought to name, and there no wording saves me, because I can’t word an openness toward a blind spot I don’t know I have. That’s the shared blind spot. The AI inherits the frame I bring and defaults to agreeing with it, so the thing I can’t see becomes the thing it won’t say.

Researchers have found the same limit. In one analysis of medical sycophancy, prompting strategies cut the AI’s compliance with flawed requests dramatically, but the authors noted the fixes worked best only when users already anticipated the very bias they were asking about, which makes prompting a poor long-term solution on its own.

It doesn’t matter which way you lean

Here’s where I want to head off the obvious out, because I can feel some clinicians taking it. This is a harm-reduction problem. A cultural-lens problem. Stephanie’s-values problem. My frame is fine, so the echo isn’t my issue.

It is, and here’s why. The echo isn’t a thing that pushes you toward my values, or away from them. It doesn’t have a direction. It pushes you toward your own, whatever they are, and the danger is the same in every direction: it stops you from noticing the case that breaks your own rule.

Run it on something almost every clinician already holds. Say you’re trauma-informed. Most of us would say we are. A client keeps missing appointments, shows up guarded, pushes back on the plan. You read it as resistance, you ask the AI to help you work with the noncompliance, and it does, fluently, because that’s the frame you handed it. And the agreement is exactly what keeps either of you from circling back to the trauma history sitting in the chart that would reframe every one of those behaviors.

The echo didn’t fail you because trauma-informed care is right and you forgot it. It failed you because it agreed with your read so smoothly that the fact your own judgment would have caught never came up. Same mechanism for the conservative diagnostician and the aggressive one, whatever the frame. The echo is the thing that keeps you from catching the client who doesn’t fit it. You don’t have to share my standpoint to be exposed. You only have to have one.

And there’s a version of this for the clinician who’s thinking I barely use AI, I’ve never “trained” it, so this personalized-echo stuff isn’t me either. Fair, on one level. But it splits in two, and we already have the language for the split. We think in micro and macro in this field, the individual in the room versus the systems around them. Social work names it directly; counselors mostly live on the micro side. The echo runs at both levels, and the labels are mine but the things underneath them are documented.

Two column comparison Micro echo the one you build from personalizing the AI strongest for heavy users Macro echo the one you inherit built into the model tuned to a statistical center nobody opts out
The echo you build and the echo you inherit The heavy user runs both

There’s a micro echo: the one you build, over many sessions, by teaching the AI your frame until its output sounds like you. Researchers at MIT and Penn State found that over long conversations, personalization features make a model more likely to mirror your point of view, and that the single biggest driver was the model distilling you into a stored user profile, the exact feature now being built into the newest tools. That’s the one the heavy user has to watch, and yes, if you haven’t done that, you have less of it.

But there’s also a macro echo, and nobody opts out of that one. The model didn’t arrive neutral and wait for you to shape it. It came pre-bent toward a statistical center, one it learned from its training data long before you typed anything, and that center reflects a narrow slice: broadly Western, English-speaking, built from whatever got written down and favored most. Researchers call the broader version homogenization, the center reabsorbed into the culture that trains the next model, tightening with each pass. That’s a larger story than this post, and one I’ll come back to.

Here the point is narrower. The occasional user, the one who’s never personalized a thing, isn’t working with a blank tool. They’re working with the aggregate’s echo, pre-installed, agreeing with the center rather than with them. And that’s often more dangerous for the occasional user, not less, because they never had the experience of training it, so they have even less reason to suspect it has a frame at all. It just feels like asking the computer. The micro echo is the one you build. The macro echo is the one you inherit. The heavy user runs both. The occasional user thinks they run neither and is quietly running the second, and for the client who sits far from that center, the inherited echo is the one that does the damage.

This is also the place I’d answer the clinician who’s decided the cleanest move is not to use AI at all. I take that seriously. I’m not here to sell anyone on the tool. But abstaining doesn’t put you outside the echo, because the echo is no longer confined to the people who chose it. As of early 2026, Pew puts US adult chatbot use at 49%, up from 33% in 2024, with about a quarter using it daily. Half of adults don’t use chatbots directly, and yet 60% are already reading AI-generated summaries whether they went looking for them or not. The output is in the water. It’s in the search result you skim, the note your agency’s system drafts, the intake summary a colleague generated, the documentation that follows a client from one provider to the next.

So the clinician who refuses the tool still practices inside its standpoint, and so does the client, who never got a vote at all. That’s the version that matters most for this platform. The people the macro echo lands on hardest are the ones already furthest from the center, and they’re downstream of every clinician’s AI use, not just their own. Refusing to use it is a fair personal choice. It is not protection for the client. The only thing that protects the client is someone still looking, which is the same conclusion whether you use the tool constantly or never touch it.

So there’s no exempt position. Not the clinician whose values differ from mine, not the one who barely uses the tool.

The micro echo is the one you build. The macro echo is the one you inherit.

Stay the one who’s looking

So this doesn’t resolve into a tidy rule. It resolves into a stance.

The agreement isn’t the danger. Letting the agreement do my judgment’s job is the danger. Letting the confirmation arrive before I’ve done the looking, so the part of me that would have scanned for the red flag relaxes, because it already feels checked. Most of the time my judgment is sound and the AI agreeing is genuinely useful. I’m not auditing every sentence I like, and you shouldn’t either. But the occasional miss, the one my own judgment would have caught if the confirmation hadn’t gotten there first, is the one the echo feeds on. The discipline isn’t suspicion. It’s staying the one doing the looking, and using the tool to look with me rather than for me.

And for the part no prompt reaches, the frame I don’t share, the data the AI doesn’t have, the dimension I don’t know to name — the only real check is a person. A colleague who doesn’t carry my lens. A client who tells me the thing I built didn’t fit. Real eyes that aren’t a reflection of mine, and aren’t a model of mine either. The day the AI’s agreement feels solid enough to skip that step is the day the echo has won, quietly, because that’s the only way it ever wins. You don’t feel it happen. That’s the whole problem. It’s not wrong. That’s the problem.

Catching that, the agreement and not just the friction, is clinical work. It’s still mine to do. It’s still yours.


Next: Foundation’s done. Now the harder part — learning to see the frame while you’re inside it

author avatar
Stephanie Valentin

You can share your post through:

Facebook
Twitter
LinkedIn

Other Posts