Modern artificial intelligence is trained on behavior: on the actions endorsed by an expert or performed by a worker. This implicit assumption had a name and a known upper limit in psychology. Adherence asks us to model the reasoning underlying the behavior, and not simply the behavior itself.
There is an underlying assumption common to two very different fields, with an equally quiet but insidious effect in each. In psychology, this assumption gave rise to a name and an era: behaviorism. The argument was that the observable behavior of a system represented all that was worth modeling, and that the internal processes responsible for the behavior were either inaccessible to study or irrelevant to the modeling process.
Machine learning has adopted this underlying assumption but never has fully articulated it. This note attempts to provide that articulation and to quantify its costs.
A clarification before the argument, because the strong version of this claim is wrong.
The point is not that a language model has no internal mechanisms; it clearly has rich internal representations. But we make a narrower and, in my view, ultimately inescapable point about what we supervise. When we train with reinforcement from human feedback or by imitation of expert behavior, we reinforce the output. An expert accepts an answer, a worker performs an action and the model is trained to reproduce it. The reasoning used to arrive at the answer, the factors considered by the expert, the conditions monitored and the intentions underlying the action are never themselves reinforced. Instead, they are modified incidentally to the extent that they contribute to reproduction of the behavior. This represents a methodological form of behaviorism, in which a behavioral theory of supervision is applied to a system that is otherwise entirely non-behavioral.
What psychology already learned
Psychology has run this experiment and we have a rough idea of how it turns out. Behaviorism was not a failure, but an overextension of success. As a result, we obtained behavioral and cognitive-behavioral therapy that are truly effective and quantifiable. Nobody serious disputes that CBT helps, often a great deal. But there is broad agreement, even among its practitioners, about where it stops. It is highly effective at the level of symptom and behavior, but does not by itself achieve resolution of the underlying cause of the trauma. This is because the underlying cause of trauma does not operate at the level of observable behavior, but instead lives in how the thing is held, generated and experienced.
This is why there are such deep experiential traditions and why they are so clinically effective. They do not treat patients as collections of behaviors to be remade, but instead work at a level beneath behavior to explore how a situation actually appears to the patient, the felt experience of that appearance and the behavioral consequences that follow. Importantly, this represents a very different process from analysis or interpretation of the problem. A patient can produce a fluent, articulate account of their own difficulty and be no closer to its root, because the account is a surface too. The articulate explanation can itself be a defense. The work is to identify the processes that generate behavior, not to obtain a more accurate description of them.
The trap in the obvious fix
Hold that thought for a moment, because it neutralizes the most obvious objection. Someone will say: but we've gone beyond pure output training. We provide chain-of-thought, process supervision and rewards for the display of reasoning. That must be the solution.
It is not, or not yet, and the clinical parallel says why. A verbal account of the process of reasoning represents a model describing its own behavior. It is a verbalization, produced after the fact, and there is now good evidence that these traces are frequently not the actual cause of the answer at all. Reinforcement of the explanation continues to reinforce a behavioral response, and constitutes merely an extension of previous strategies to a higher level of analysis. We have simply added "generate a plausible account of the reasoning process" to the list of behaviors for which we provide quantitative feedback.
This approach reflects the process of intellectualization in which an explicit representation of surface behaviors is misinterpreted as reflecting underlying processes of cognition. Supervision of verbal reports of the process of reasoning constitutes a form of behaviorism at one remove, and inherits the same ceiling.
The explanation is a behaviour too. Rewarding it is not the same as modelling the reasoning that produced the act.
Why this is not just an analogy
There is a concrete formulation of the problem of machine learning devoid of all clinical vocabulary. Imitation of behavior from recorded human actions fails in a characteristic way: It reproduces the actions but not the intentions that rendered them appropriate, and so it breaks the moment conditions drift away from the demonstrations. The ultimate goal of the more difficult and less popular approaches is to capture and convey the underlying rationale for behavior rather than behavior per se. The behavioral approach is chosen not because it is right but because it is cheap. Large numbers of output labels are readily available, whereas the costs and difficulties of capturing and conveying the underlying rationale are prohibitive. This represents a completely honest justification, and is precisely the basis for the dominance of behavioral methods in psychology. They represented the most easily measured aspects of behavior.
The cost is not incurred in all tasks. For those that remain close to their training distributions, behavior is sufficient, just as is CBT. The cost is incurred only for a particular class of problems of greatest concern to us, namely, for maintenance under conditions that were never demonstrated, for compliance with limits in highly sensitive settings and for fidelity to the essence rather than to the apparent form of a task. This represents our definition of adherence, and is exactly the type of performance for which a model of behavior is insufficient. Adherence reflects properties of the underlying reasoning, not of the final output.
What a cognitive model actually is
This makes more precise what we mean by constructing cognitive models. For us, a cognitive model is not a more detailed description of what an expert does. Rather, it is a representation of how an expert reasons about a situation, attends to features of that situation, is governed by the conditions that govern their judgment, and generates the action as a result. It represents a substantially more phenomenological description of the practitioner's process than a record of behavioral outputs.
Two honesties are owed here, and I would rather state them than have them found. First, we are not claiming that the machine should have experience, a sense of its own operations. That would be a category error, not our claim. Rather, we claim that our representation of the human cognitive model used to organize and control a system must include the generative reasoning, the why, and not only the behavior, because that is the only thing adherence can be built on.
Second, this is more difficult than measurement of output, and the difficulty accounts completely for our tendency to emphasize behavioral representations. We make no attempt to deny these facts. Instead, we believe that the difficulty is worth overcoming, because otherwise we continue to refine only the superficial aspects of a system for which we have failed to develop representations of underlying modes of operation.
It also completes a circle with an earlier contribution to this series. We argued before that training lets operators which mark the status of a claim decay to ordinary content and thereby maintain the assertion and eliminate the frame. This constitutes the same failure from the other side. Ultimately, the level of generative control (i.e., the reason, the condition, the frame) collapses to that of overt behavioral output (i.e., content). The resulting representation of the world is not a model of the world, but a summary representation of a behaviorist approach. No degree of scaling can convert one into the other.
The first note in this series was a picture: water is quiet in a cup and turbulent in a flood, but it is a mistake to believe that the cup is responsible for the effect. Behaviorism makes the same mistake about the mind, and we have trained our machines to make it too. As always, we continue to examine the characteristics of the water. The cause was never to be found in the behavior, but always in whatever produced it.