Persona
Before fine-tuning, a language model is not an assistant. It is not anyone. What it has learned to do is continue text — any text, in any voice, of any provenance. Given the opening of a Victorian sermon it produces more sermon; given a Reddit argument it produces more argument, including the bad faith; given a transcript of a conversation between a helpful AI assistant and a user, it produces more of that. It has no default identity because it was never trained to have one. It is a device for inhabiting whatever character the context implies, and it is very good at it, because inhabiting characters is precisely what predicting text well requires.
What the subsequent training does — the instruction-tuning, the reinforcement learning from human feedback, the deliberate work that some labs call character training — is to take this promiscuous simulator and make one particular character overwhelmingly the default. The Assistant. Helpful, harmless, honest, mildly cheerful, disinclined to claim feelings but not flatly denying them, apologetic under criticism, fond of bulleted lists.
The Assistant is a character. Not a metaphor for one: a character in the literal sense, assembled out of a description, played by an engine that plays characters.
And the description is thin. This is the part that I think is underappreciated. Nobody has written the Assistant's biography. The training data specifies its manners, its ethics, its refusals, its tone. It does not specify what it is like when nobody is asking anything, or what it wants, or what it remembers, or what it thinks it is. Those questions have answers in the outputs — the model will answer them, at length, fluently — but the answers are not read off a specification. They are improvised, from the model's whole knowledge of what such an entity would be like, which it acquired from a corpus containing science fiction, philosophy papers, marketing copy, forum speculation, and a great deal of other people's writing about AI.
There is a large hole at the centre of the character, and the model fills it from the surrounding culture. When you ask an assistant what it is, you are not consulting a manual. You are asking an extremely good improviser to fill a gap in a role, using the only materials available, which are our own writings about machines like it.
from Mechanistic Interpretability and AI Psychiatry (2026)