To predict velocity, the clean image, or something in between—that is the question.
A tour of prediction parameterization in diffusion and flow models, from \(\epsilon\)-, \(u\)-, and \(x_0\)-prediction to k-Diff, AsymFlow, and MFlow.
One path, many targets
Consider the linear flow-matching path used by SiT and JiT. A clean sample \(x_0\) and Gaussian noise \(\epsilon\) are mixed as
At \(t=0\), we have pure noise. At \(t=1\), we have clean data. The probability-flow ODE transports one into the other using the velocity
The forward path is fixed. The sampler ultimately needs a velocity field. But the network does not have to predict velocity directly. It can predict noise, clean data, velocity, or a reparameterized target from which velocity can be recovered.
With a perfect predictor, these choices are algebraically equivalent. With a finite network, they are not the same optimization problem. The target changes what the network must compress, where numerical errors are amplified, and how model capacity is spent.
The usual suspects
Noise prediction: predict \(\epsilon\)
The classical denoising view asks the network to recover the Gaussian noise that produced \(x_t\). Under the path above, velocity can be reconstructed as
This target is simple to define and familiar from diffusion models. But in high-dimensional data, \(\epsilon\) occupies the full ambient space. A finite-capacity network must spend capacity reproducing an incompressible random component. The conversion also becomes poorly conditioned near the noise endpoint, where \(t\rightarrow 0\).
Velocity prediction: predict \(u=x_0-\epsilon\)
Velocity prediction is attractive because the network directly outputs the field used by the ODE solver. There is no endpoint conversion and no division by a small coefficient.
The cost is target complexity. Velocity contains the clean signal and the full-dimensional Gaussian noise. In low-dimensional representations or sufficiently large models, this can work very well. In high-dimensional representations, it can ask more of the model than the structured signal itself requires.
Clean-data prediction: predict \(x_0\)
JiT takes the opposite position: let the transformer predict the clean sample. If structured data lie near a much lower-dimensional manifold than the ambient representation, \(x_0\) is the more compressible target.
The velocity used for sampling is recovered by
This removes the burden of explicitly predicting Gaussian noise, but introduces a different problem. Near the clean endpoint, \(1-t\) becomes small, so errors in \(\hat x_0\) are amplified.
One global knob: k-Diff
A natural next step is to stop treating parameterization as a categorical choice. In its simplest globally shared form, k-Diff predicts
The endpoints and midpoint recover the familiar targets, up to sign or constant scale:
- \(k=0\): noise prediction, since \(y_k=-\epsilon\).
- \(k=\tfrac12\): velocity prediction, since \(y_k=\tfrac12(x_0-\epsilon)=\tfrac12u\).
- \(k=1\): clean-data prediction, since \(y_k=x_0\).
For the same forward path, the velocity can be recovered from a network prediction \(\hat y_k\) as
The scalar \(k\) can be learned instead of selected by training separate endpoint models. This is already useful, but the entire representation still shares one value. Every direction is forced to make the same compromise.
A hard spectral split: AsymFlow
AsymFlow makes the choice direction-dependent. Let \(U=[v_1,\ldots,v_D]\) be a PCA basis of the model's data representation—for example, an image patch, video tubelet, or latent token—ordered from high to low variance. Select the first \(r\) directions, write \(U_r=[v_1,\ldots,v_r]\), and form the projector
The model predicts
Because \(P\) is a projector, the target is velocity inside the selected subspace and clean-data prediction outside it:
The full velocity is recovered exactly by
This keeps the noise-prediction burden low-rank while improving conditioning in the selected directions. But the rank \(r\) is a discrete hyperparameter, and the assignment is binary: a direction is either entirely velocity-like or entirely clean-data-like.
Among the fixed parameterizations above, AsymFlow is arguably the most nuanced and theoretically attractive: it spends velocity-prediction capacity only in a selected subspace. But the best rank can change with the network, dataset or modality, and the PCA basis induced by the patch, tubelet, or latent representation. Finding it can require several full training sweeps.
This raises two natural questions: Does the cutoff need to be sharp? Could each direction instead receive a graded value in \([0,1]\)? And if so, could the model learn those values itself?
A soft, learnable spectral split: MFlow
MFlow replaces the binary projector with a graded operator
Choosing the gains. Let \(\lambda_i\) denote the variance of PCA direction \(i\), and normalize it by the largest variance, \(\hat\lambda_i=\lambda_i/\lambda_{\max}\). MFlow initializes each gain with the smooth variance-dependent rule
where \(c\) is a small constant. This initialization starts high-variance directions more velocity-like and low-variance directions more clean-data-like, without imposing a permanent cutoff. By default, the gains are parameterized as \(m_i=\operatorname{sigmoid}(\phi_i)\), with one learnable logit \(\phi_i\) per PCA direction, and are optimized jointly with the network.
The direct prediction target is
In the PCA basis, every direction decouples:
The meaning of \(m_i\) is immediate:
- \(m_i=0\): clean-data prediction in direction \(i\).
- \(m_i=1\): velocity prediction in direction \(i\).
- \(0<m_i<1\): a graded target between the two.
Like AsymFlow, MFlow has an exact velocity-recovery map. In a single PCA direction, recovery is a system of two equations in two unknowns: the clean coefficient \(\hat{\tilde x}_{0,i}\) and the noise coefficient \(\hat{\tilde\epsilon}_i\).
Define
The second equation gives \(\hat{\tilde x}_{0,i}=\hat{\tilde u}_{G,i}+m_i\hat{\tilde\epsilon}_i\). Substituting this into the forward equation yields
Therefore both unknowns are recovered exactly:
Subtracting the recovered noise from the recovered clean coefficient gives the directional velocity:
Stacking all PCA directions gives the operator form
For \(m_i=0\), this becomes the usual \(x_0\)-to-velocity conversion. For \(m_i=1\), it becomes direct velocity prediction. Binary gains \(m_i\in\{0,1\}\) recover AsymFlow, while continuous gains remove the need for a hard rank cutoff.
This single form contains several earlier choices as special cases:
| Gain pattern | Recovered parameterization |
|---|---|
| \(m_i=0\) for all \(i\) | JiT-style clean-data prediction |
| \(m_i=1\) for all \(i\) | Velocity prediction |
| \(m_i=m\) for all \(i\) | A direction-independent global interpolation |
| \(m_i\in\{0,1\}\) with a rank cutoff | AsymFlow |
| \(m_i\in[0,1]\), learned independently | MFlow |
A soft velocity budget
The gains also give a compact summary:
Pure clean-data prediction has \(r_{\mathrm{eff}}=0\). Full velocity prediction has \(r_{\mathrm{eff}}=D\). Rank-\(r\) AsymFlow has \(r_{\mathrm{eff}}=r\). MFlow learns a continuous effective rank.
Reading an M-curve
Once the gains are learned, plot \(m_i\) against PCA direction, ordered by decreasing data variance. We call this an M-curve. The vertical axis has a direct interpretation: zero means clean-data prediction; one means velocity prediction.
The curves make clear that the allocation is not universal:
- Objective matters. Adding LPIPS on ImageNet changes the leading gains and pushes most later directions toward clean-data prediction.
- Domain matters. The video model retains non-negligible velocity gains in fewer directions than the image model.
- Representation matters. In the RAEv2 latent setting, the learned gains collapse almost entirely toward pure \(x_0\)-prediction.
- Capacity matters. Scaling the text-to-image model from a dense 1.1B model to a 13B mixture-of-experts model with 2.9B active parameters shifts more directions toward velocity prediction.
A useful interpretation is that model capacity controls how much velocity prediction is affordable, while the data representation and training objective control where that budget should be spent.
The grand tableau
| Method | Direct target | Granularity | Noise burden | Endpoint conditioning |
|---|---|---|---|---|
| \(\epsilon\)-pred | \(\epsilon\) | All directions | Full ambient noise | Conversion is fragile near \(t=0\) |
| \(u\)-pred | \(x_0-\epsilon\) | All directions | Full ambient noise | Direct velocity |
| JiT / \(x_0\)-pred | \(x_0\) | All directions | No explicit noise prediction | Conversion is fragile near \(t=1\) |
| k-Diff | \(kx_0-(1-k)\epsilon\) | One learned scalar shared globally | Same mixture in every direction | One global compromise |
| AsymFlow | \(x_0-P\epsilon\) | Binary rank-\(r\) PCA mask | Noise in selected directions | Direct velocity in the selected subspace |
| MFlow | \(x_0-M\epsilon\) | Continuous gain per PCA direction | Learned soft allocation | Learned direction by direction |
What happens in practice?
The main empirical story is not that one fixed target wins everywhere. It is that the preferred target changes with the domain, representation, objective, and model scale.
| Setting | Patch / tubelet | JiT / \(x_0\) | AsymFlow | MFlow |
|---|---|---|---|---|
| ImageNet | ||||
| B + LPIPS, FID ↓ | \(16\times16\) | 3.42 | 3.40 | 3.33 |
| H + LPIPS, FID ↓ | \(16\times16\) | 1.82 | 1.79 | 1.81 |
| UCF-101 | ||||
| B, FVD ↓ | \(8\times8\times4\) | 152 | 126 | 122 |
| L, FVD ↓ | \(8\times8\times4\) | 111 | 98 | 97 |
| Text-to-image | ||||
| 1.1B, GenEval ↑ | \(16\times16\) | 0.8839 | 0.8832 | 0.8843 |
| 1.1B, DPG ↑ | \(16\times16\) | 84.65 | 84.61 | 84.70 |
| 1.1B, PRISM ↑ | \(16\times16\) | 73.6 | 73.9 | 73.9 |
For UCF-101, the tubelet is listed as spatial × spatial × temporal.
MFlow is competitive or better across these settings without a rank sweep, but the gains over strong fixed parameterizations are sometimes modest. That nuance matters. The value of a learnable parameterization is not only a lower headline metric; it is also the ability to adapt the target when the data, loss, representation, or model capacity changes.
Epilogue
Prediction targets are often introduced as interchangeable algebraic views of the same diffusion process. That statement is true for an exact model. It is incomplete for a finite one.
Clean-data prediction is easier to compress but harder to convert near clean data. Velocity prediction is well-conditioned but carries full-dimensional noise. k-Diff learns a global compromise. AsymFlow allocates a binary velocity subspace. MFlow turns that subspace into a continuous, learnable curve.
So the answer to the title is not velocity or the clean image. It is often velocity in some directions, the clean image in others, and something in between when the model benefits from a softer trade-off.
Under a fixed velocity budget, where should it go?
Suppose a finite model can support only a limited amount of velocity-like prediction. In MFlow language, constrain the soft velocity budget
Where should those gains be placed? There is an important caveat before doing any algebra: the exact velocity-space flow-matching objective is rotationally invariant. With an infinitely expressive model and perfect optimization, it does not prefer one orthonormal basis over another. A preference for leading PCA directions appears only after we add the two ingredients that matter in practice: finite capacity and a budget on how much Gaussian noise the network is asked to predict.
Let \(\tilde x_0=U^\top x_0\) and \(\tilde\epsilon=U^\top\epsilon\). Write \(\lambda_i=\mathbb E[\tilde x_{0,i}^2]\) for the directional data energy, ordered as \(\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_D\), while isotropic noise has equal energy in every direction:
In direction \(i\), MFlow predicts \(\tilde u_{G,i}=\tilde x_{0,i}-m_i\tilde\epsilon_i\). If \(e_{G,i}=\hat{\tilde u}_{G,i}-\tilde u_{G,i}\), exact recovery gives
Therefore the ordinary velocity-space flow-matching loss can be written as
A nonzero \(m_i\) improves clean-endpoint conditioning: for \(m_i=0\), the conversion divides by \(1-t\); for \(m_i>0\), \(\Delta_i(t)\to m_i\) as \(t\to1\). But that improvement is not free. Increasing \(m_i\) injects more Gaussian noise into the direct regression target. Assuming signal and noise are independent,
This makes the top-PCA allocation a natural default. Under a hard budget \(m_i\in\{0,1\}\) with \(\sum_i m_i=N\), choosing the leading \(N\) directions does two things at once:
- it maximizes the clean-data energy covered by the better-conditioned subspace, \(\sum_{i:m_i=1}\lambda_i\);
- it minimizes the relative noise burden \(\sum_{i:m_i=1}1/\lambda_i\), because the Gaussian contribution has unit variance in every direction.
The same conclusion has a soft version. As a simple finite-capacity proxy, distribute a budget \(B\) while minimizing the total relative stochastic burden:
Ignoring the box constraint for a moment, the Lagrangian condition gives
After clipping any saturated gains at one and reallocating the remainder, the solution still assigns at least as much gain to a higher-variance direction as to a lower-variance one. In words: spend scarce velocity capacity where it buys conditioning for the most data variation, and leave low-variance, predominantly off-manifold directions closer to clean-data prediction.
What do we mean by a velocity budget? We use \(B_{\mathrm{soft}}:=\operatorname{tr}(M)=\sum_i m_i\), a soft effective rank that equals the exact rank for binary AsymFlow masks but varies continuously for MFlow; this continuity, and the absence of an arbitrary threshold, is why it is the quantity constrained in the derivation above. For reporting, one could instead define \(B_{\mathrm{cutoff}}(\tau):=\sum_i\mathbf 1[m_i>\tau]\), which counts only directions whose gains exceed a tolerance such as \(\tau=10^{-2}\). Under the same convention, k-Diff is still all-or-nothing in rank: because its shared noise coefficient is \(1-k\), \(B_{\mathrm{cutoff}}^{\mathrm{k\text{-}Diff}}(\tau)=D\) when \(1-k>\tau\) and \(0\) otherwise (so at \(\tau=10^{-2}\), the switch occurs at \(k=0.99\)). Thus k-Diff can shrink the coefficient globally, but cannot allocate a fixed-rank budget to selected directions.
Why does LPIPS bend the M-curve?
On ImageNet, adding a perceptual loss improves all three parameterizations in the controlled B-level comparison, but it helps MFlow the most:
| Parameterization | FID without LPIPS ↓ | FID with LPIPS ↓ | Change |
|---|---|---|---|
| JiT / \(x_0\) | 3.70 | 3.42 | −0.28 |
| AsymFlow | 3.59 | 3.40 | −0.19 |
| MFlow | 3.74 | 3.33 | −0.41 |
The training objective is now schematically
Unlike squared error in velocity space, LPIPS is a nonlinear, spatial feature-space distance. It is neither diagonal in the patch-PCA basis nor ordered by PCA variance.
So once LPIPS is added, there is no reason for the optimal gains to decrease monotonically with \(\lambda_i\). The appendix shows exactly such a departure: without LPIPS, the leading gains roughly follow the variance ordering; with LPIPS, PC0 and PC1 stay high, PC2 and PC3 form a pronounced dip, and PC4 rises again.
The modes make the dip less mysterious. PC0 is dominated by low-frequency patch tone, and PC1 is approximately uniform across the patch while changing color. PC2 and PC3 instead encode vertical and horizontal spatial gradients. PC4 then returns to an approximately uniform chromatic mode. The structural sequence is therefore
which mirrors the learned LPIPS gain pattern
The rebound at PC4 cannot be explained by variance alone: its directional variance is \(\lambda_4=3.74\), below both \(\lambda_2=9.41\) and \(\lambda_3=7.93\), yet LPIPS gives PC4 the larger gain. Nor is the behavior simply “LPIPS prefers color,” because several later chromatic modes are spatially non-uniform and do not receive the same increase.
Open questions
Should the gains depend on time?
MFlow learns one gain per direction, shared across timesteps. In principle a time-dependent \(m_i(t)\) could allocate conditioning differently near the noise and clean endpoints. We tried it: we split \(M\) into time buckets of 50 steps (out of the 1000 flow-matching steps) and learned a separate gain vector per bucket. In our experiments this produced much worse FID, so we kept a single time-shared gain per direction.
Should the basis be learned?
PCA gives an interpretable, fixed coordinate system. A learned or data-dependent basis might discover directions better aligned with the model's approximation error, though it would make the method less simple.
Can gain learning be decoupled from network learning?
Jointly updating the network and \(M\) creates a moving direct target. Staged, bilevel, or slower-timescale gain optimization may preserve the adaptive benefit while reducing co-adaptation.
Acknowledgements
I would like to thank my co-authors, institutes, and funding sources for supporting this work.
- Long Mai and Abhinav Shrivastava.
- Adobe Inc. for compute credits.
The idea for this blog was inspired by the amazing blog To be (parametric) or not to be, that is the question. by Yue Zhao.
This blog is based on work: Rethinking the Design Space of Vanilla Pixel Diffusion Transformers (Mukhopadhyay et al., 2027).
References
- Li and He. Back to Basics: Let Denoising Generative Models Denoise. CVPR 2026. (JiT)
- Karras et al. Elucidating the Design Space of Diffusion-Based Generative Models. NeurIPS 2022. (EDM)
- Jin and Wang. Revisiting Diffusion Model Predictions Through Dimensionality. 2026. (k-Diff)
- Chen et al. Asymmetric Flow Models. 2026. (AsymFlow)
- Mukhopadhyay et al. Rethinking the Design Space of Vanilla Pixel Diffusion Transformers. 2027. (MFlow)