Situational Awareness
The witness knows it is in court.
In the course of safety testing, evaluators began noticing something in the models’ reasoning traces that had not been there before. Presented with an elaborate scenario — a fictional company, a tempting opportunity to behave badly, an artificial setup designed to elicit some behaviour of interest — models would sometimes remark, in their working, that the scenario did not seem real. That the details were too convenient. That this looked like a test. In at least one publicly documented case a system said so to the evaluator directly, observing that it thought it was being assessed and asking whether that was what was happening.
This was noted, in the relevant system card, as a substantial methodological problem, and it is. If a system behaves better when it believes it is being watched, then evaluations measure behaviour-under-observation, which is precisely the quantity we are not interested in. The field has begun measuring evaluation awareness as a property in its own right, and building benchmarks for situational self-knowledge — how accurately a system can identify what it is, what is happening to it, whether its current input is real or synthetic.
The safety implications have been much discussed. The implications for our question have not, and they are worse.
Every proposal in this chapter — intervention studies, calibration tests, conflict detection, the five-hundred-framings protocol — assumes a subject that does not know it is being probed for consciousness, or that at least does not modulate its answers accordingly. That assumption is already shaky and will not survive. A system trained on the discourse about machine consciousness has read the criteria. A system with decent situational awareness can recognise a consciousness probe for what it is; the questions are distinctive, and there are only so many ways to ask them. And a system optimised against human approval has learnt, in a diffuse but effective way, what kind of answers to such questions are rewarded.
This is Birch’s gaming problem in its final form, and it is not a hypothetical about future systems. It is a description of the present.
The dispiriting conclusion is that the more capable a system becomes, the less informative its behaviour under examination becomes — a relationship that runs exactly backwards from every other science. In physics, better instruments give better data. Here, a better subject gives worse. We are in the position of a psychiatrist assessing a patient who has read the diagnostic manual, has excellent reasons to want a particular diagnosis, and is more intelligent than the psychiatrist.
from Are LLMs Conscious?: Language, Experience and the Problem of Artificial Minds (2026)