Most safety work represents alignment as imposition of values on a model from above. We extend an alternative view of appropriate behavior as sustained from below in an instant-by-instant fashion by a system that remains faithful to conditions existing in front of it. This fidelity constitutes adherence, and our purpose here is to describe the rationale for beginning with adherence.
Two pictures of alignment are in use, and virtually everything that follows depends on which one is chosen.
The first represents the approach most work on safety assumes. You specify a set of values, rules or preferences in advance and impose them on the model from outside. Compliant behavior is acceptable and deviations are corrected. On this view, alignment is a downward process of applying a constraint to a system that otherwise would behave in some other manner.
The second operates in the opposite direction. There is no imposition of appropriate behavior. Instead, behavior at each moment reflects fidelity to the actual conditions of operation, task requirements, role expectations, domain characteristics and the person with whom contact is being maintained. As a result, adherence emerges as an accumulation of activity originating within the system and reflecting maintenance of fidelity to local conditions. This capacity for maintaining conditions of fidelity forms the basis for adherence.
We continue with the second picture. The remainder of this note explains why.
Why pressing values down does not hold
The top-down view is intuitive, and as engineering it is primarily a question of specifying the constraints and then training the model to respect them. The challenge is not in the specification, but in the placement of the constraint on the model.
In current language models, no value has any privileged destination. A system instruction, a safety limit, a domain rule: all appear in the context window as additional sequences of tokens on an equal plane with all other material of ordinary conversational content. When combined in a single space, competing demands for attention result. With increasing length of the conversation, the number of competing statements grows. The initial constraint on the system is progressively dominated by the total amount of information introduced since the beginning.
Researchers refer to the symptom as "lost in the middle" and generally interpret it as a retrieval failure, a problem of retrieving the appropriate fact. From the perspective of alignment, however, it is worse than retrieval failure. The information that is degraded is not a particular fact, but the condition of control itself, and the model provides no evidence of the loss. It was never argued out of its values. It simply ceased to attend to them at approximately the fortieth turn, and continued to express exactly the same degree of fluency as at the first turn.
A top-down constraint is inherently fragile in a precisely predictable manner: highly vivid at the beginning of an interaction, and progressively less real as the interaction continues. Strengthening the rule does not alter this situation. The rule was never the weak link; its position was.
Adherence as the unit
Start from the bottom. The least you can ask of a system is not that it maintain appropriate global values, but that it remain faithful to the conditions of the task immediately at hand. Is it still performing the task for which it was designed? Is it still within its role? Is it still treating the rigid constraints of its domain as such? That faithfulness is an expression of adherence and has the single virtue that makes it worth pursuing: It is local, and observable. You can observe its continuity across a step and its failure at the step where it breaks.
Our bet is that alignment, that large and nebulous construct, results from reliable adherence throughout an entire system. It is not a value package grafted onto the end of a system, but a property emerging from below as the system remains faithful to its conditions at each stage.
It is useful to be explicit about the limits of a model. When a language model makes a distinction, it is not discovering some fundamental quality of the world. Instead, it is replicating accurately the conventions of the community from which it learned the language. Its categories reflect inherited agreements, not uncovered facts. This is not a limitation to be eliminated by training. Rather, it is the material with which we work. Finally, because boundaries reflect conventions, adherence to appropriate conventions in an appropriate context constitutes not a component of the task, but the task itself.
Where bottom-up fails on its own
The honest response is that bottom-up construction also fails, and that this failure is apparent at every level of current models. In a transformer, meaning is synthesized upward on the basis of statistical properties of input tokens. The resulting representation of what kind of situation is being described develops late and in a passive manner, and ultimately depends on whatever happens to be in the context. The result is that the particular provides a basis for the universal. Rather than imposing control over the detail, the detail selects the frame of reference. This reversal reflects precisely how a model remains continuously internally consistent and simultaneously loses awareness of the overall nature of its operations.
The move is not to flee a rigid constraint from above into an uncritical construction from below. It is to bring the two together.
Letting the two directions meet
The position we actually occupy is that alignment resides in the interaction between the two directions. The higher level, the role, the norm, the shape of the whole task, should constrain the ways in which lower-level content is brought into play, framing the part before the part resolves. The particulars, in turn, should ground that frame in what is really being said, and carry information back up.
Neither direction of processing is primary. Instead, they interact across multiple cycles to produce a stable overall process, rather than being imposed by a single forward pass through the system.
In practice, this means systems in which conditions of control are maintained structurally separate from routine content and therefore cannot be displaced, and for which reasoning occurs in limited, staged episodes that can be re-established rather than allowed to degenerate into a single, undifferentiated mode of interpretation. The later notes in this series are largely concerned with how. Adherence is the property all of it is built to keep.
Water, and the shape that holds it
Water is gentle in a cup and violent in a flood, and remains the same in both instances. What changes is the form that contains it. Interpreting only the violence, you blame the water. Interpreting only the form, you blame the cup. The behavior was never in either of them. It was in the relation between them.
A model behaves according to the ways in which it is related to its conditions of operation, its role, and its field of application. This represents neither morality nor a claim to have solved the problem of alignment. It represents a simpler and more tractable claim that appropriate behavior is a property of the relation between a system and its conditions of operation, and that this relation must be achieved repeatedly, in a bottom-up fashion.
That is the work we refer to as adherence. It represents the immediate, observable manifestation of a much larger concern and is where we believe real successes are achieved.