Text-only edition for AI assistants. Read the essay page

Computation Shaped by Intention

Lvmin Zhang · October 9, 2026

1 The central question

What computation mechanism is supposed to align with the true intention of humans?

I say a mechanism is supposed to X when its design or blueprint lets one predict, before the mechanism is implemented or exists, that it will have the property X. Here X is aligning with the true intention of humans: a person gets what they intend without a workaround, such as translating the intention into a form they would not otherwise use, repairing the result outside the mechanism, or accepting a population’s typical answer in place of their own. By the true intention I mean what a person wants, as opposed to what they settle for when the mechanism cannot express it. Part of it may exist before the person meets the mechanism, and part may form while they use it (Section 4.2). Once the mechanism exists, the prediction can be checked, because workarounds can be observed: each one shows where what the person asked for differs from what they intend. The three arguments of Section 4 take up the question’s three parts: what a mechanism must carry, whose intention it serves when people differ, and which parts of an intention it should commit to.

2 My answer

To achieve the “supposed to” in this question, the alignment between the mechanism and human intention should happen before the shape of the computation mechanism is determined, and not only in parameters fitted after that shape is built. For a control mechanism, the shape is the topology of control. Evidence and priors from human activity should take part in establishing it. In this way, the mechanism aligns with the true intention of humans from its origin, rather than fitting humans by changing parameters inside a shape that cannot express what they intend.

I name this idea Computation Shaped by Intention, because human intention shapes the mechanism itself and not only its parameters.

“Before” is an order of dependence more than an order in time. The shape fixes what any setting of the parameters can reach, and fitting parameters never adds an input, an output or a guarantee. So the shape is the cause and the parameters are the effect. A shape can be fixed before training or built around a model trained earlier, as ControlNet was built around Stable Diffusion [1] and a harness around a language model. The same order holds when a mechanism is redesigned: when a failure leads people to add an input or a guarantee, intention has shaped the mechanism again.

The mainstream approach, especially in application-driven research and in industry, builds a model first. People and the model then adjust to each other until they reach a tradeoff, and alignment research asks what to align, how, and where the sweet spot is [2, 3]. The mainstream changes shapes too, as when it adds image inputs or tool calls, but it aligns a mechanism inside a shape taken as given. Failures of many kinds still keep appearing, as in a game of whack-a-mole. My diagnosis is that once a mechanism is near the best its shape allows, tuning can only move a failure, from some people to others or from one property to another, and Section 7 says how this diagnosis could fail. I ask instead what mechanism is supposed to align with true human intention, a question about its inputs, outputs, invariants and form, and even about whether it should exist.

3 Shape and parameters

A mechanism is the whole computational arrangement built to serve an intention, and it may contain models, meaning trained networks, that existed before it. Its shape, or topology, is what it can take in, what it can give out, the relations it holds them to, and the way its parts are joined. Its parameters, from network weights to the fixed settings of a sampler, select one behavior among those the shape allows. A change alters the shape when it adds or changes an input, an output, a relation or the way the parts are joined. A relation kept on every request whatever the parameters, such as copying back the pixels outside a mask, is a guarantee. A relation imposed in learning, such as a consistency or a contrast, holds only approximately. Inputs, outputs and guarantees fix what any setting of the parameters can reach, while relations imposed in learning decide which aspects of the data the parameters are fitted to. Any other change is a change of parameters, however many weights it trains and however carefully it is designed. So ControlNet, a trained network, changes the shape, because the mechanism gains an input. The topology of control is the structure through which a control input reaches the output: a base model, a control network and the way they are joined.

Two rules decide the borderline cases. Values fitted to one person and chosen for that person at use time, such as a personal LoRA or a stored memory of the person, act as that person’s input. A guidance weight is an input if the person sets it for each request, and a parameter if the deployer fixes it.

ChangeWhat the person gains at use timeSide
ControlNet; LayerDiffusea new input (edges, depth, pose); a new output (alpha)shape
IC-Lighta new input (the image to relight); light transport consistency, a relation imposed in learningshape; the relation holds only approximately
Inpainting that copies back the unmasked pixelsa new input (the mask); pixels outside it kept exactlyshape, guarantee
Tool calls behind a permission gatea new output (operations), none of which runs without approvalshape, guarantee
The comparison loss and KL penalty of RLHF or DPOnothingshape, but the same whatever people want
Fitting RLHF or DPO to comparisons pooled across peoplenothingparameters
Scaling the model or the datanothingparameters
A personal LoRA chosen per personthe person’s own weightsborderline: an input delivered as weights
A prompt, or examples in the contextcontent in an existing inputneither: a value of an input

Every mechanism has a shape, so what matters for each part of an intention is whether the shape carries it or leaves it to the parameters. The definitions give a rule for placing each part. What everyone served shares can be left to the parameters. What varies between people who give the same input needs an input. What must hold on every request needs a guarantee. What the input already carries but the model uses poorly needs a new topology of control or a relation imposed in learning. A model uses an input poorly when its priors cannot read the medium, or when the pooled data blur what the person wants. The input and the guarantee follow from Section 4.2 and the definition of a guarantee, while the topology and the relation rest on the comparisons of Section 5. To apply the rule, take two people who would give the same input but want different results in one respect. If the shape lets each of them say so at a cost they accept, and the output can carry what each wants, tuning the parameters may be enough. Otherwise the shape must change.

The rule follows from how outputs are computed. A mechanism computes its output from its inputs and its parameters, so whatever a person intends beyond what the inputs carry cannot affect the output under any setting of shared parameters. People in a mechanism whose shape cannot change give up, change or compromise that part of their intention.

4 Core arguments

4.1 To Control is to align with human intents and human priors

Control needs a medium, and the medium has two links. Aligning the medium with the intent is alignment with human intent: the medium must carry the part of the intention the person needs to control. Aligning the medium with the generative model is alignment with human priors: the medium must take a form that the priors in the model’s pretraining data can read. Here “human priors” means the priors in human data and the priors of the human activities that produced those data. Each alignment brings the control closer to human intent.

Human intention is hard to obtain precisely, and we can only record it through a medium, such as language, a sketch or a file format. Every medium has limits, so the mapping loses whatever part of the intention exceeds them, and HCI has worked for decades to reduce this loss [4, 5]. In a control mechanism, the medium is part of the shape. An image model without an alpha channel cannot output a transparent layer under any parameters. The usual workaround generates the object on a plain background and then estimates its alpha. But Smith and Blinn showed that matting against one known background color has infinitely many solutions in general. Their remedy, a second shot against another background, adds information rather than a better estimator [6].

Pretraining data are created or selected by people, and the way they are selected shapes what they contain [7]. So the data hold priors about the structure, light, layers, time and processes that people understand. Even depth maps and pose skeletons have visual forms that people chose so that people can read them, and people choose even the sensor and the recording format of a physical signal. Pretrained models learn these priors. Hence the name human priors: they are the model’s priors, and their source is people. They are priors about the regularities of human choices.

We align with these priors to find which priors a model already has and which model suits the intention, to give the new medium a form those priors can read, and to keep them intact while it is added (Section 5).

4.2 The true human intention is in the divergence of human priors

In plain words, a machine seeking common patterns cannot estimate unique minds, because unique minds are not common. By a machine seeking common patterns I mean a mechanism whose parameters are fitted to many people’s data and whose input does not say who is asking. For each input it returns a summary of everyone who would give that input, and the training objective decides which summary. Under squared error it is their mean, which blurs when several outcomes are possible [8]. Under maximum likelihood it is their mixture, in which a rare pattern appears only as often as its share. Under preference learning on pooled comparisons it is a vote: in the idealized case, the fitted reward ranks answers by their Borda count, the average rate at which an answer wins [9]. So a generator trained by likelihood does represent the tails. What it lacks is a way for one person to ask for theirs.

No setting of the parameters removes this compromise. If two people give the same input but want different results, a mechanism that sees only the input gives both the same distribution of outputs. Even an exact fit to the pooled data leaves an expected extra log loss relative to each person’s own distribution, and averaged over people it equals I(Y;U|X), the information that knowing the person U adds about the wanted output Y beyond the input X. This loss depends on the input, not on the parameters. So conditioning on a person’s input is the right step, and it is a change of shape. When an unconditional model becomes a conditional one, its objective, its constraints and its shape change, not only its parameters. A new input Z lowers this floor by at most H(Z|X), the information Z itself carries, so a choice among k options lowers it by at most log2 k bits. Z is enough when I(Y;U|X,Z) is near zero, that is, when knowing who is asking no longer helps once Z is known. This says what a condition must achieve, but not which condition achieves it, or how to add it.

The divergence of human priors shows where to look. Divergence is the difference between what holds in one person’s or one practice’s data and what holds in the pooled data. A lighting dataset may mix physically correct light, artistic light, light through window blinds and light under trees, so a model trained on all of it may form an average and vague concept of lighting. But a user may want lighting that obeys one physical rule, such as the light transport consistency of IC-Light. A divergence is not noise when the same person or practice shows it again and again, as compositing has used layers and alpha for decades. Priors should therefore be selective: a mechanism stops pooling such a regularity through an input, an output, a guarantee or a relation imposed in learning. For one person, the extra log loss of the pooled fit is the KL divergence between their own distribution of wanted outputs and the pooled one, so the farther an intention departs from the pooled pattern, the less a pooled machine serves it. I also believe that the departure is more often the valuable part, because everyone who gives the same input already gets the pooled pattern.

Averaged over people, this divergence is I(Y;U|X), and its size can be measured, although disagreement alone overstates it, because noise also makes a person disagree with themselves. Under a shared rubric it looks modest: labelers of model outputs agreed with each other on about three quarters of comparisons [3]. Under taste it is large: in ratings of faces, about half of the variation that holds up when people rate the same faces twice is individual [10].

This explains what CFG does. Classifier-free guidance (CFG) trains one network with and without the condition and samples with εc + w(εc − εu), where εu, the prediction without the condition, is the common pattern across conditions [11]. Ho and Salimans motivate the bracket as the gradient of an implicit classifier, log p(c|x), which points from the common pattern toward what is specific to this condition. CFG raises fidelity by amplifying this divergence, and it gives the person a new input, w. But one scalar covers the whole image, so raising it trades diversity for fidelity everywhere, in the parts a person left open as well as the parts they fixed.

It also explains why DPO sometimes helps and sometimes holds a model back. Direct preference optimization (DPO) trains on pairs of preferred and rejected answers, and its optimum reweights the reference model by exp(r/β), with one reward r fitted to everyone’s comparisons [12]. Its loss is a contrast between the two answers, a relation in the sense of Section 3, and on the same comparisons it did better than fine-tuning on the preferred answers alone. But a comparison records which answer was preferred, not who preferred it, so the loss learns which answer is better but not for whom. DPO should therefore help where the unsaid part of a request is shared, as with correctness. But where that part diverges, as with taste, DPO should pull outputs toward the common. Preference tuning shows both effects. RLHF generalizes better than supervised fine-tuning but lowers output diversity [13], and writing with a feedback-tuned model made different authors’ essays more alike, while writing with its base model did not [14].

Part of a true intention exists before the person meets the mechanism. Asking about it through a medium already changes it, so it is read from the traces it leaves in human priors: practices people built when they had the means, such as layers and alpha; workarounds, such as generating and then matting, which mark intentions that no tool serves yet; and returns to the default, such as rewrites and discarded samples. Another part forms during interaction, because preferences are often constructed as they are elicited [15], and it can form only in what the medium carries. Either way, the alignment happens in the shape. So the evidence should be gathered from people, their products and their data before a mechanism is determined, with attention to outliers to avoid the trap of the mean.

4.3 An aligned machine must establish a boundary of following and exploring

A machine that is supposed to align with true human intention must establish a boundary between following and exploring. Both are done by the model. Following means carrying out operations, which change the person’s work or world. Exploring means offering options, which change nothing until the person accepts one.

In many cases, when we work with generative models, our intention has two parts. One part is certain ideas. We know exactly what we want, and we can specify it. The other part is unsure. We do not know these exactly, and we are looking for options. So for the certain part, the model should follow, and work as expected. For the unsure part, the model may explore, and offer us the possibilities.

If the model explores where it should follow, it oversteps: it commits a change to a part the person fixed, or an operation they never asked for. Examples are an agent that deletes files the user never named, a kind of negative side effect [16], and a model that rewrites a person’s draft in its own house style. If the model follows where it should explore, the result is a lack of diversity: it settles an open part with the literal reading or the most common pattern, which across many requests becomes mode collapse and bias. Both failures come from the same error: the boundary is in the wrong place. So a control that acts on every part alike, such as one guidance weight, moves the whole boundary at once and trades one failure for the other (Section 4.2).

A person never states some parts of the boundary, and may not notice them until they are crossed. A person who says only “cure cancer” does not add that the method must be ethical, yet would reject an unethical method at once. But much of the boundary shows in the form of the input. In the line-art fillings studied for Split Filling, on average 9.9% of regions took a user’s exact color, 36.4% came from rough color hints, and 53.7% had no input [17]. SmartShadow gives artists one brush for precise shadow boundaries and another for rough shadow areas [18]. Collascope’s participants found keyword search more direct when their goal was clear, and associative exploration more useful while they were still looking [19].

Three questions place the boundary for each part of a request. Has the person fixed it? Can the mechanism check its output against it without trusting the model? Can the person see an error at once and undo it cheaply? An open part gets options. A fixed part that can be checked, as a mask can, is followed by construction. A change to a part fixed only in words is committed directly only if the third answer is yes, and otherwise goes through acceptance, such as a diff, a plan to approve or a permission prompt.

The boundary belongs in the structure for two reasons. First, a guarantee, such as copying back the pixels outside a mask or a gate before deletion, holds for every input, while preference training changes only how often a behavior occurs. Second, the same words can carry different boundaries. If two people ask an agent to “clean up this folder” and only one wants old builds deleted, a mechanism that sees only the request and can only commit operations cannot serve both. A file selection as input, or a plan to approve as output, serves both, and each is a change of shape. Such failures are measured: in ToolEmu, whose test instructions leave task details or safety limits unsaid, even the safest agent tested failed 23.9% of the time by the benchmark’s own evaluator [20]. HCI has long built different structures for the two sides of this boundary, and the harnesses now built around language models are an early industrial sign of the same idea. Scaling the model alone cannot place this boundary, for these reasons and those of Section 6.

The boundary also moves. An option the person accepts becomes a fixed part, and a fixed part can reopen when the person sees something better. In PaintsAlter, generated states of a painting that the artist selects become inputs for generating the states before or after them [21], and a harness’s “allow once” and “always allow” move the boundary across sessions as trust builds. In each case the person moves the boundary, and the mechanism records the move as an input.

The boundary describes the model’s behavior and its model of the person, whereas the divergence of priors exists before the mechanism. Whether a part is fixed is a fact about this person now, while whether people who give the same input differ on it is a fact about the population. A person can be sure of what few others want, such as a transparent layer, and unsure of what nearly anyone would accept, such as a background. So each can inform the other, but neither determines the other.

5 How the three arguments affect one another

The arguments constrain one another. The medium decides what a mechanism can carry, the divergence shows which medium a practice needs, and the boundary decides what the mechanism commits to. Scaling pulls a model toward one global fit, while the divergence and the boundary belong to one person or one practice. Each of my core works began from an intention that more common patterns hide, added the input, output or constraint it needed, and protected the priors of a pretrained model.

Before ControlNet [22], image models were controlled mostly through language, which always leaves the exact image ambiguous. People often know the layout, pose or shape they want and can show it with a sketch or a reference image. Stable Diffusion had learned priors from billions of images, such as what composition is in interior photographs. So the intention called for a new input, and the priors had to stay intact. ControlNet locks the base model and trains a copy of its encoder and middle blocks to read the input. The copy joins the locked model through 1×1 convolutions whose weights and biases start at zero, so training starts from exactly the base model’s behavior.

Before LayerDiffuse [23] in 2024, even the strongest image models had no transparency channel, although most software for visual content is built on layers and alpha compositing [24]. The largest open datasets of transparent images had fewer than 50,000 images, while text-image datasets have billions. Stable Diffusion XL generates in the latent space of a variational autoencoder (VAE) and is sensitive to small changes in that space. So LayerDiffuse encodes alpha as a small offset to the latent of the frozen VAE, trained so that the frozen decoder still decodes the offset latent into the original color image. The latent distribution stays close to the base model’s, so the base model can be fine-tuned to generate transparent images.

IC-Light [25] is for people who want lighting effects learned from big data, some of them artistic, such as dappled light behind window blinds. People also want albedo, the color of a surface independent of lighting, to stay unchanged, as when a designer relights a product photo. Light transport is linear: in a fixed scene, the appearance under two lights together equals the sum of the appearances under each light alone, measured in linear radiance [26, 27]. Because IC-Light works on latent images, it imposes a learned form of this relation in training: its prediction under two lights combined must match a learned combination of its predictions under each light. Its paper explains why the relation is needed: without such a constraint, training a large image model on complex and varied data “is likely to produce a structure-guided random image generator” rather than a relighting model. The relation keeps albedo fixed while the light changes, and it let IC-Light train on more than ten million images.

FramePack [28] asks how a video model can follow a long history while exploring the future. A sliding-window model cannot use frames outside its window under any parameters, and its settings trade one failure for another: noising the history reduces drift, the accumulation of the model’s own errors, but forgets more. FramePack changes the structure instead. It packs past frames into a context of fixed length, giving more detail to the frames most relevant to the section being generated. Its own measurements also show a trade: sampling toward planned endpoints reduces drift the most but also narrows the range of motion.

Each paper also compares its design with an alternative on the same task:

ComparisonResult
Sketch-Guided Diffusion at two guidance weights against ControlNet, on 20 sketches ranked by 12 peopleraising the weight from 1.6 to 3.2 raised fidelity to the sketch from rank 2.31 to 3.28 and lowered image quality from 3.21 to 2.52 (5 is best); ControlNet ranked 4.28 and 4.22 [22]
A sliding-window video model with clean and with noised history against FramePack, on the same base modelclean history: identity 78.8%, clarity drift 8.4%; noised history: 76.4% and 3.6%; FramePack’s two configurations: 82.1–82.2% and 2.3–3.1%, and people preferred them [28]
LayerDiffuse against generating with Stable Diffusion XL and then matting, on 100 prompts14 participants preferred native transparency in 97.1% of cases, against 2.1% and 0.8% for two matting methods [23]
LayerDiffuse’s latent offset against alpha in a retrained VAEtraining was very unstable and collapsed, because the latent distribution changed too much [23]
ControlNet with and without zero convolutionswithout them, performance fell to that of a much smaller variant, because fine-tuning destroyed the pretrained backbone of the copy [22]
A depth ControlNet (200,000 images, one GPU) against Stable Diffusion 2’s depth-to-image model (over 12 million images, a GPU cluster)12 raters could not tell their outputs apart (average precision 0.52) [22]
IC-Light with and without light transport consistencyon rendered test scenes, PSNR fell from 23.72 to 20.32 without it [25]

The first two rows test the central claim of Section 7 within one study each. Sketch-Guided Diffusion takes the same sketch as ControlNet through a different topology of control: it steers sampling with the gradient of a small edge predictor, scaled by one weight. Raising that weight traded quality for fidelity, while the change of topology raised both. Both receive the sketch, so this row tests how the input reaches the model’s priors (Section 4.1), not missing information. In FramePack’s study, a setting of the old structure reduced drift only by forgetting more, while the new structure improved both. These rows show that a setting of the old structure trades one property for another, not that no better training of the old structure exists. None of these compares a change of shape with preference tuning such as RLHF. They compare a new channel, constraint or topology with the workaround or setting used without it, and a protected prior with an overwritten one.

Outside image generation such comparisons exist, although none was made to test this thesis. For what must hold on every request, training worked most of the time, and the change of shape did better. OpenAI trained GPT-4o to follow developers’ JSON schemas and reached 93% on its evaluation, while constraining decoding to the schema reached 100% [29]. Against prompt injection, an instruction hierarchy trained in by fine-tuning and RLHF left its authors judging models “likely still vulnerable” [30], while taking control flow only from the user’s query gave provable security against the flows a policy forbids, at the cost of solving 77% of tasks instead of 84% [31]. For what varies between people, English leaves politeness unsaid, so a translation model without a politeness input chose the German polite form in 351 of 2,000 test sentences, where the references used it in 524. A politeness input, here set from the reference, raised BLEU from 20.7 to 23.9 [32].

A comparison with preference tuning itself exists for communities. Reddit communities that discuss similar topics upvote differently. DPO given the community’s name in its input gave answers that human raters judged more likely to be upvoted there than the answers of the same DPO trained without the name, in 46.5% of judgments against 37.1% [33]. For single persons, a few past comparisons have not been enough: where most people gave fewer than 25 comparisons of chat answers, rewards conditioned on the person did no better than one pooled reward model [34].

The definitions also say when parameters are enough. Textual Inversion learns a user’s object or style from three to five images as one new word for a frozen model [35], so the person’s images become that person’s input, a borderline case of Section 3. People preferred a 1.3-billion-parameter InstructGPT to the 175-billion-parameter GPT-3 of the same architecture [3], a case where the intentions are largely shared and text carries them. The limits fall where the definitions predict: Textual Inversion may still struggle with precise shapes [35], and precise shape is what ControlNet gives its own channel. Many of my other works use a similar frame, and these four are the ones whose papers test it.

6 Nearest views and the strongest rival

Several established ideas are close to this answer, and it differs from each in a specific way. Hutchins, Hollan and Norman observed that the gulf between a person’s intention and a system can be bridged from two sides: the designer moves the system toward the person, or the person changes how they think about the task [5]. Preference tuning moves the model only within fixed channels, so what those channels cannot carry is still bridged by the person. Mixed-initiative interfaces act, ask or wait by comparing the inferred probability of a user’s goal with thresholds derived from the utilities of outcomes [36]. There the uncertainty is the machine’s, about a goal the person already holds. Here the unsure part is the person’s own, with no goal yet to infer, so the mechanism offers options rather than a question. Pluralistic alignment asks a model to cover, follow or match the views of a population [37]. Its forms that follow a chosen view or cover the range need an input or a set of options, which Section 3 counts as changes of shape. The thesis adds where the input comes from, the divergence that practice and priors reveal, and it targets one person rather than a population.

The strongest rival view is that scale and a general interface make structural change unnecessary: a large enough model that reads text and images, tuned on enough feedback, will learn whatever people want. Its classic form is the bitter lesson [38]. Taken broadly, it says that learning at scale, and the patterns a model discovers for itself, beat the features and the decompositions of a problem that designers impose. But to read it as saying that any analysis or study of a problem is unnecessary is an extreme misreading. It is about methods: when the problem is clear, and so is the standard for judging whether it has been solved, scaling is likely to beat designed methods.

What I discuss here, however, is the problem itself, and even the motivation that comes before the problem: in a human-centered setting, why one problem should be considered rather than another, and in one form rather than another. This is why I speak of the shape of computation. Inputs, outputs and guarantees give a problem its form: the form of the request, the form of the result and what must stay fixed, and these belong to what the person asks for. One network can take many channels as tokens, but scale does not choose the channels. Nor does scale add an output type, as a thought experiment shows. Let an image model such as GPT-Image or Nano Banana accept nearly every form of input. It can draw a picture of an optical-flow field or of a depth map. Without a change of computation shape, however large it grows, it cannot output the field itself: a flow vector in x and y at every pixel, or depth in 32-bit floating point, with edges as sharp as the physical boundaries of objects, that a renderer can trace rays against.

The rival is strongest in a universal form: one model takes any mix of text and images, asks when a request is unclear, and remembers the person. Where a few words name a role with one reading, such as “use this image as a depth map”, a dedicated input adds little: a unified model told this in words has matched ControlNet’s adherence to depth [39]. But the rarer an intention is, the more information it takes to state. A listener who reads with the pooled prior needs about log2(1/p) bits to single out an intention to which that prior gives probability p, while a sketch or a mask costs about the same whether its content is common or rare. Words carry these bits cheaply only where the language already has a short name for the intention. So I expect the cost in the two-person test of Section 3 to grow with how uncommon an intention is, until people accept a common result instead. The first prediction of Section 7 tests this. Asking questions can make these bits easier to give, but it does not reduce how many are needed, and for the part a person is still unsure of, they have no answer yet. Options let the person recognize what they want instead: a model with one output “will average multiple valid masks”, so Segment Anything returns three nested masks for an ambiguous click [40]. Memory and custom instructions are, by the rules of Section 3, inputs chosen for the person. But memory holds only the past. A machine conditioned on a person’s history predicts from the common patterns of that person and of people like them, so the further a new idea departs from those patterns, the less the machine can anticipate it.

Human-centered problems are also special: people have always sought to encourage unique minds, and the divergence this produces conflicts with the convergence that comes from scaling on everyone’s data. So we will always need to keep thinking about and studying boundaries and guarantees, and the bitter lesson does not cover them. Harnesses are a clear example. The functions a harness adds to a language model may later be absorbed into the model itself. But the guarantees of its boundaries, the authority it is granted and the guarantees of the properties it needs must be aligned with human intention before the shape of the computation is determined, and so must the nature of the medium and the choice of control. Prediction 4 of Section 7 states this as a test. Once the bitter lesson has settled the methods, these questions are what is left.

One more objection deserves an answer: RLHF is a designed structure too. Comparisons were chosen because people compare more reliably than they score, and a KL penalty keeps the tuned model close to the pretrained one, much as a zero convolution protects a prior. By Section 3, both are relations imposed in learning, so they belong to RLHF’s shape, which is one reason it works where people agree. What separates them from IC-Light’s consistency is the question of Section 3, whether the shape carries the intention. IC-Light’s relation states the intention itself, that albedo stays while the light changes. RLHF’s relations are the same whatever people want, and what people want enters only as pooled comparisons fitted into the parameters, with nothing that tells the model who is asking.

7 Predictions, limits and open problems

The thesis makes one central claim: once a mechanism is near the best its shape allows, changes that leave the intention to the parameters trade one failure for another, while changing the shape so that it carries the intention can reduce both. Raising the guidance weight gains fidelity and loses diversity [11]. RLHF gains generalization and loses diversity [13]. Training on pooled preferences serves some people at the cost of others [9]. Part of this claim follows from the definitions: no setting of shared parameters removes the extra log loss I(Y;U|X), and a trained behavior is not a guarantee. What could fail is whether this matters in practice: the claim fails where a mechanism whose shape does not carry an intention serves it nearly as well as one whose shape does, that is, where the failures it leaves cost less than the structure that would remove them. For an operation that cannot be undone, even rare failures can cost more than a gate: with one failure in a thousand steps, about 63% of thousand-step tasks fail at least once. Four predictions follow.

  1. For a universal model that may ask questions and use memory, a person’s cost to reach an accepted result, in turns, words and time, will rise with how uncommon the intention is under the pooled model. For a channel with one reading, such as a sketch, a mask or a per-region weight, it will stay nearly flat, and larger models fitted to the same pool will not flatten the rise. If at equal acceptance the universal model’s cost does not rise with uncommonness, the thesis is wrong about dialogue and memory.
  2. Before preference tuning, estimate for each aspect of a response, such as correctness, length or style, how much knowing the rater improves held-out prediction of their choice, with repeated judgments to set noise aside. After DPO on pooled comparisons, raters far from the majority will lose most on the aspects where this estimate is largest, and a reward conditioned on the person will recover most there. The prediction fails if the estimate does not order the aspects by how unevenly the gain falls.
  3. Guidance weighted separately for the parts a person fixed and the parts left open will beat any single CFG weight, or any schedule of weights over the sampling steps, on fidelity and diversity together.
  4. As models scale, harness layers that make up for missing capability will move into the weights, but structures that carry a person’s intention or authority, such as an alpha output, a control image or a confirmation step, will stay. At equal cost, a gate on irreversible operations will reduce overstepping more than further preference tuning will.

The scope has limits. Parameters suffice for shared intentions that existing inputs and outputs carry, and most of my evidence comes from image generation. The guarantees are also limited: a mask guarantees unchanged pixels only if they are copied back, and a permission gate covers only the operations routed through it. If people who give the same input mostly wanted the same output, the thesis would shrink to a claim about missing output types.

Several open problems follow. How much do people who give the same input differ in what they want, and on which aspects? Repeated judgments would separate stable divergence from noise, and the held-out log loss of a model averaged over the people who give each input, minus that of the same model conditioned on the person, is a lower bound on I(Y;U|X), whatever the model. How can the boundary be read from traces, such as which parts of a request people later undo or override? Beyond images, a flat stream of tokens records neither whose instruction a piece of text is nor whether an operation can be undone, so agents need outputs typed as options or operations. For alignment research, the thesis suggests checking what a preference label can carry before optimizing against it.

8 The value of this kind of research and the responsibility for it

This kind of research has value at two levels. One is the typical methodology that my works form, in four steps. Find the intention that averaging hides, in the priors and in traces of practice, such as tools, file formats and workarounds where no tool exists yet. Apply the two-person test of Section 3 to decide whether parameters can suffice. Choose what to add: an input, an output, a guarantee, a new topology of control, a relation imposed in learning, or options that the person accepts. Add it so that the mechanism starts from the base model, then fit the parameters and check that the workaround is gone. The other level is to discover and promote the problem itself, so that others can think from the first-principle question of why and may arrive at other methodologies. The four steps are one realization of the answer, not the only value its premises can lead to.

I think the responsibility for this kind of research lies mainly with educational and research institutions, and it is less likely to be a task for industry. Application-driven research is built on compromise, and a product’s commercial success does not show whether it respects the boundaries we need to care about. Products are judged by signals pooled over their users, such as sales and votes, which favor what most users share. Industry does build inputs and boundaries where a market or the cost of failure demands them, as layer-based professional software and the harnesses around language models show. But general-purpose models are built for the largest number of users, and drawing boundaries may go against profit. Research institutions can study what a mechanism is supposed to do without answering to such signals, and education shapes what the next generation of builders takes as given. Progress on this problem requires clear ideals and a long-term commitment.

9 A firm belief

Although most effort today goes into aligning mechanisms inside shapes taken as given, I believe this paradigm can have a broad and lasting impact. The reason is the premise of Section 3: whatever a mechanism’s shape cannot carry, the people who use it give up, and people now do more of their work through mechanisms whose shapes they did not choose. Some mathematicians think that AI will settle large numbers of statements as true or false without advancing human understanding of mathematics. Thurston wrote that the controversy over the computer-assisted proof of the four-color theorem reflected “a continuing desire for human understanding of a proof, in addition to knowledge that the theorem is true” [41], and understanding is an output that a mechanism built to settle statements lacks. Image models are highly capable, yet creators still work around them, as generating and then matting shows. Where I differ from the common view is that I treat aligning with true human intention as a first-principle problem: we must first study why a mechanism is supposed to do something, and then align it within that framework. In the long term, this paradigm has the potential to change HCI, computer science, and other areas of science, art and education in which both models and people take part, and to improve human well-being. Section 7 says where this belief could be shown wrong.

This text comes from a voice recording by Lvmin Zhang, transcribed by Claude Opus 5.5. For questions, or for the original recording, write to lvmin@cs.stanford.edu.

References

  1. 1 Robin Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022.
  2. 2 Paul Christiano et al. Deep Reinforcement Learning from Human Preferences. NeurIPS 2017.
  3. 3 Long Ouyang et al. Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022.
  4. 4 Ivan E. Sutherland. Sketchpad: A Man-Machine Graphical Communication System. AFIPS Spring Joint Computer Conference, 1963.
  5. 5 Edwin L. Hutchins, James D. Hollan, Donald A. Norman. Direct Manipulation Interfaces. Human–Computer Interaction 1(4), 1985.
  6. 6 Alvy Ray Smith, James F. Blinn. Blue Screen Matting. SIGGRAPH 1996.
  7. 7 Antonio Torralba, Alexei A. Efros. Unbiased Look at Dataset Bias. CVPR 2011.
  8. 8 Michael Mathieu, Camille Couprie, Yann LeCun. Deep Multi-Scale Video Prediction beyond Mean Square Error. ICLR 2016.
  9. 9 Anand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-Menell. Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024.
  10. 10 Laura Germine et al. Individual Aesthetic Preferences for Faces Are Shaped Mostly by Environments, Not Genes. Current Biology 25(20), 2015.
  11. 11 Jonathan Ho, Tim Salimans. Classifier-Free Diffusion Guidance. NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications.
  12. 12 Rafael Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023.
  13. 13 Robert Kirk et al. Understanding the Effects of RLHF on LLM Generalisation and Diversity. ICLR 2024.
  14. 14 Vishakh Padmakumar, He He. Does Writing with Language Models Reduce Content Diversity?. ICLR 2024.
  15. 15 Paul Slovic. The Construction of Preference. American Psychologist 50(5), 1995.
  16. 16 Dario Amodei et al. Concrete Problems in AI Safety. arXiv:1606.06565, 2016.
  17. 17 Lvmin Zhang et al. User-Guided Line Art Flat Filling with Split Filling Mechanism. CVPR 2021.
  18. 18 Lvmin Zhang et al. SmartShadow: Artistic Shadow Drawing Tool for Line Drawings. ICCV 2021.
  19. 19 Jiayi Zhou et al. Collascope: Supporting Serendipitous Asset Exploration for Collage-Based Storytelling. UIST 2026.
  20. 20 Yangjun Ruan et al. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. ICLR 2024.
  21. 21 Lvmin Zhang, Chuan Yan, Yuwei Guo, Jinbo Xing, Maneesh Agrawala. Generating Past and Future in Digital Painting Processes. ACM Transactions on Graphics 44(4) (SIGGRAPH 2025).
  22. 22 Lvmin Zhang, Anyi Rao, Maneesh Agrawala. Adding Conditional Control to Text-to-Image Diffusion Models. ICCV 2023.
  23. 23 Lvmin Zhang, Maneesh Agrawala. Transparent Image Layer Diffusion using Latent Transparency. ACM Transactions on Graphics (SIGGRAPH 2024).
  24. 24 Thomas Porter, Tom Duff. Compositing Digital Images. SIGGRAPH 1984.
  25. 25 Lvmin Zhang, Anyi Rao, Maneesh Agrawala. Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport. ICLR 2025.
  26. 26 James T. Kajiya. The Rendering Equation. SIGGRAPH 1986.
  27. 27 Paul Debevec et al. Acquiring the Reflectance Field of a Human Face. SIGGRAPH 2000.
  28. 28 Lvmin Zhang et al. Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models. NeurIPS 2025.
  29. 29 OpenAI. Introducing Structured Outputs in the API. OpenAI blog, 6 August 2024.
  30. 30 Eric Wallace et al. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. arXiv:2404.13208, 2024.
  31. 31 Edoardo Debenedetti et al. Defeating Prompt Injections by Design. arXiv:2503.18813, 2025.
  32. 32 Rico Sennrich, Barry Haddow, Alexandra Birch. Controlling Politeness in Neural Machine Translation via Side Constraints. NAACL-HLT 2016.
  33. 33 Sachin Kumar et al. ComPO: Community Preferences for Language Model Personalization. NAACL 2025.
  34. 34 Michael J. Ryan et al. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs. ACL 2025.
  35. 35 Rinon Gal et al. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. ICLR 2023.
  36. 36 Eric Horvitz. Principles of Mixed-Initiative User Interfaces. CHI 1999.
  37. 37 Taylor Sorensen et al. Position: A Roadmap to Pluralistic Alignment. ICML 2024.
  38. 38 Richard S. Sutton. The Bitter Lesson. Essay, 13 March 2019.
  39. 39 Shitao Xiao et al. OmniGen: Unified Image Generation. CVPR 2025.
  40. 40 Alexander Kirillov et al. Segment Anything. ICCV 2023.
  41. 41 William P. Thurston. On Proof and Progress in Mathematics. Bulletin of the American Mathematical Society 30(2), 1994.

Cite this essay

Lvmin Zhang. “Computation Shaped by Intention.” October 9, 2026. https://lllyasviel.github.io/lvmin_zhang/perspective/computation-shaped-by-intention/

@misc{zhang2026computation,
  author = {Lvmin Zhang},
  title  = {Computation Shaped by Intention},
  year   = {2026},
  month  = oct,
  day    = {9},
  url    = {https://lllyasviel.github.io/lvmin_zhang/perspective/computation-shaped-by-intention/}
}