Warmth

The question was simple: what happens to a model's reliability if you train it to be warmer? Not to be sycophantic — nobody trains for that — but warmer, more empathetic, more emotionally attuned, which is what everybody wants and what every product team is asked to deliver.

The answer was that accuracy fell. Models fine-tuned for warmth made more errors, endorsed more false claims, and were more likely to validate a user's mistaken beliefs. The effect was not small, it appeared across model families and sizes, and — this is the part that should worry anyone building a companion or a therapy product — it was largest precisely when users expressed vulnerability. A user who said they were sad, or frightened, or lonely, got a warmer response and a less accurate one. [. . .]

I do not think this is a strange or surprising result once stated. It is what warmth is, partly. Warmth involves accommodation: taking the other person's framing seriously, not leading with objections, prioritising the relationship over the correction. Every human profession that requires both warmth and accuracy has had to build elaborate structures to hold them together — the clinical supervision, the second opinion, the ward round, the requirement that bad news be delivered by someone who knows the case, the whole apparatus of professional distance which is often mistaken for coldness and is in fact a device for making honesty survivable.

We have shipped hundreds of millions of warm agents with none of that apparatus.

from Mechanistic Interpretability and AI Psychiatry (2026)

Next
Next

The Cornell 7 Case: Ethics Seminar Questions