September 2026  ·  Aniruddha Mahapatra and Soumik Mukhopadhyay

To predict velocity, the clean image, or something in between—that is the question.

A tour of prediction parameterization in diffusion and flow models, from \(\epsilon\)-, \(u\)-, and \(x_0\)-prediction to k-Diff, AsymFlow, and MFlow.

One path, many targets

Consider the linear flow-matching path used by SiT and JiT. A clean sample \(x_0\) and Gaussian noise \(\epsilon\) are mixed as

\[x_t = t x_0 + (1-t)\epsilon, \qquad t\in[0,1].\]

At \(t=0\), we have pure noise. At \(t=1\), we have clean data. The probability-flow ODE transports one into the other using the velocity

\[u = \frac{d x_t}{dt} = x_0-\epsilon.\]

The forward path is fixed. The sampler ultimately needs a velocity field. But the network does not have to predict velocity directly. It can predict noise, clean data, velocity, or a reparameterized target from which velocity can be recovered.

With a perfect predictor, these choices are algebraically equivalent. With a finite network, they are not the same optimization problem. The target changes what the network must compress, where numerical errors are amplified, and how model capacity is spent.

The usual suspects

Noise prediction: predict \(\epsilon\)

The classical denoising view asks the network to recover the Gaussian noise that produced \(x_t\). Under the path above, velocity can be reconstructed as

\[\hat u = \frac{x_t-\hat\epsilon}{t}.\]

This target is simple to define and familiar from diffusion models. But in high-dimensional data, \(\epsilon\) occupies the full ambient space. A finite-capacity network must spend capacity reproducing an incompressible random component. The conversion also becomes poorly conditioned near the noise endpoint, where \(t\rightarrow 0\).

Velocity prediction: predict \(u=x_0-\epsilon\)

Velocity prediction is attractive because the network directly outputs the field used by the ODE solver. There is no endpoint conversion and no division by a small coefficient.

The cost is target complexity. Velocity contains the clean signal and the full-dimensional Gaussian noise. In low-dimensional representations or sufficiently large models, this can work very well. In high-dimensional representations, it can ask more of the model than the structured signal itself requires.

Clean-data prediction: predict \(x_0\)

JiT takes the opposite position: let the transformer predict the clean sample. If structured data lie near a much lower-dimensional manifold than the ambient representation, \(x_0\) is the more compressible target.

The velocity used for sampling is recovered by

\[\hat u = \frac{\hat x_0-x_t}{1-t}.\]

This removes the burden of explicitly predicting Gaussian noise, but introduces a different problem. Near the clean endpoint, \(1-t\) becomes small, so errors in \(\hat x_0\) are amplified.

The core trade-off is not simply which target is correct? It is where should we pay the price: target complexity or numerical conditioning?

One global knob: k-Diff

A natural next step is to stop treating parameterization as a categorical choice. In its simplest globally shared form, k-Diff predicts

\[y_k = kx_0-(1-k)\epsilon, \qquad k\in[0,1].\]

The endpoints and midpoint recover the familiar targets, up to sign or constant scale:

For the same forward path, the velocity can be recovered from a network prediction \(\hat y_k\) as

\[\hat u=\frac{\hat y_k-(2k-1)x_t}{k(1-t)+(1-k)t}.\]

The scalar \(k\) can be learned instead of selected by training separate endpoint models. This is already useful, but the entire representation still shares one value. Every direction is forced to make the same compromise.

The limitation: a global coefficient can answer how much of one target to use, but not where in the representation to use it.

A hard spectral split: AsymFlow

AsymFlow makes the choice direction-dependent. Let \(U=[v_1,\ldots,v_D]\) be a PCA basis of the model's data representation—for example, an image patch, video tubelet, or latent token—ordered from high to low variance. Select the first \(r\) directions, write \(U_r=[v_1,\ldots,v_r]\), and form the projector

\[P=U_rU_r^\top.\]

The model predicts

\[u_A=x_0-P\epsilon.\]

Because \(P\) is a projector, the target is velocity inside the selected subspace and clean-data prediction outside it:

\[P u_A=P(x_0-\epsilon)=Pu,\qquad(I-P)u_A=(I-P)x_0.\]

The full velocity is recovered exactly by

\[\hat u=P\hat u_A+(I-P)\frac{\hat u_A-x_t}{1-t}.\]

This keeps the noise-prediction burden low-rank while improving conditioning in the selected directions. But the rank \(r\) is a discrete hyperparameter, and the assignment is binary: a direction is either entirely velocity-like or entirely clean-data-like.

Among the fixed parameterizations above, AsymFlow is arguably the most nuanced and theoretically attractive: it spends velocity-prediction capacity only in a selected subspace. But the best rank can change with the network, dataset or modality, and the PCA basis induced by the patch, tubelet, or latent representation. Finding it can require several full training sweeps.

This raises two natural questions: Does the cutoff need to be sharp? Could each direction instead receive a graded value in \([0,1]\)? And if so, could the model learn those values itself?

A soft, learnable spectral split: MFlow

MFlow replaces the binary projector with a graded operator

\[M=U\,\mathrm{diag}(m_1,\ldots,m_D)U^\top,\qquad m_i\in[0,1].\]

Choosing the gains. Let \(\lambda_i\) denote the variance of PCA direction \(i\), and normalize it by the largest variance, \(\hat\lambda_i=\lambda_i/\lambda_{\max}\). MFlow initializes each gain with the smooth variance-dependent rule

\[ m_i^{(0)}=\frac{\hat\lambda_i}{\hat\lambda_i+c}, \qquad \hat\lambda_i=\frac{\lambda_i}{\lambda_{\max}}, \]

where \(c\) is a small constant. This initialization starts high-variance directions more velocity-like and low-variance directions more clean-data-like, without imposing a permanent cutoff. By default, the gains are parameterized as \(m_i=\operatorname{sigmoid}(\phi_i)\), with one learnable logit \(\phi_i\) per PCA direction, and are optimized jointly with the network.

The direct prediction target is

\[u_G=x_0-M\epsilon.\]

In the PCA basis, every direction decouples:

\[\tilde u_{G,i}=\tilde x_{0,i}-m_i\tilde\epsilon_i.\]

The meaning of \(m_i\) is immediate:

Like AsymFlow, MFlow has an exact velocity-recovery map. In a single PCA direction, recovery is a system of two equations in two unknowns: the clean coefficient \(\hat{\tilde x}_{0,i}\) and the noise coefficient \(\hat{\tilde\epsilon}_i\).

\[ \begin{aligned} \tilde x_{t,i} &= t\,\hat{\tilde x}_{0,i}+(1-t)\,\hat{\tilde\epsilon}_i,\\ \hat{\tilde u}_{G,i} &= \hat{\tilde x}_{0,i}-m_i\,\hat{\tilde\epsilon}_i. \end{aligned} \]

Define

\[\Delta_i(t)=(1-t)+tm_i.\]

The second equation gives \(\hat{\tilde x}_{0,i}=\hat{\tilde u}_{G,i}+m_i\hat{\tilde\epsilon}_i\). Substituting this into the forward equation yields

\[ \tilde x_{t,i} =t\,\hat{\tilde u}_{G,i}+\Delta_i(t)\,\hat{\tilde\epsilon}_i. \]

Therefore both unknowns are recovered exactly:

\[ \begin{aligned} \hat{\tilde\epsilon}_i &=\frac{\tilde x_{t,i}-t\hat{\tilde u}_{G,i}}{\Delta_i(t)},\\[4pt] \hat{\tilde x}_{0,i} &=\frac{(1-t)\hat{\tilde u}_{G,i}+m_i\tilde x_{t,i}}{\Delta_i(t)}. \end{aligned} \]

Subtracting the recovered noise from the recovered clean coefficient gives the directional velocity:

\[ \hat{\tilde u}_i =\hat{\tilde x}_{0,i}-\hat{\tilde\epsilon}_i =\frac{\hat{\tilde u}_{G,i}-(1-m_i)\tilde x_{t,i}}{\Delta_i(t)}. \]

Stacking all PCA directions gives the operator form

\[ \hat u =\big[tM+(1-t)I\big]^{-1} \big[\hat u_G-(I-M)x_t\big]. \]

For \(m_i=0\), this becomes the usual \(x_0\)-to-velocity conversion. For \(m_i=1\), it becomes direct velocity prediction. Binary gains \(m_i\in\{0,1\}\) recover AsymFlow, while continuous gains remove the need for a hard rank cutoff.

This single form contains several earlier choices as special cases:

Gain patternRecovered parameterization
\(m_i=0\) for all \(i\)JiT-style clean-data prediction
\(m_i=1\) for all \(i\)Velocity prediction
\(m_i=m\) for all \(i\)A direction-independent global interpolation
\(m_i\in\{0,1\}\) with a rank cutoffAsymFlow
\(m_i\in[0,1]\), learned independentlyMFlow
Figure 1. A schematic parameterization atlas. Bar height indicates how velocity-like each direction is. Asym8 is a rank-8 AsymFlow mask; the MFlow row is illustrative, not a measured curve.

A soft velocity budget

The gains also give a compact summary:

\[r_{\mathrm{eff}}=\mathrm{tr}(M)=\sum_i m_i.\]

Pure clean-data prediction has \(r_{\mathrm{eff}}=0\). Full velocity prediction has \(r_{\mathrm{eff}}=D\). Rank-\(r\) AsymFlow has \(r_{\mathrm{eff}}=r\). MFlow learns a continuous effective rank.

Reading an M-curve

Once the gains are learned, plot \(m_i\) against PCA direction, ordered by decreasing data variance. We call this an M-curve. The vertical axis has a direct interpretation: zero means clean-data prediction; one means velocity prediction.

Four learned MFlow gain curves for ImageNet, UCF-101 video, RAEv2, and text-to-image models.
Figure 2. Learned M-curves across ImageNet, UCF-101 video, an RAEv2 latent representation, and two text-to-image model scales. Directions are ordered by decreasing PCA variance; \(m_i=1\) is velocity-like and \(m_i=0\) is clean-data-like.
Use velocity-like prediction in as many directions as the model can accurately support, and keep the remaining directions clean-data-like.

The curves make clear that the allocation is not universal:

A useful interpretation is that model capacity controls how much velocity prediction is affordable, while the data representation and training objective control where that budget should be spent.

The grand tableau

MethodDirect targetGranularityNoise burdenEndpoint conditioning
\(\epsilon\)-pred\(\epsilon\)All directionsFull ambient noiseConversion is fragile near \(t=0\)
\(u\)-pred\(x_0-\epsilon\)All directionsFull ambient noiseDirect velocity
JiT / \(x_0\)-pred\(x_0\)All directionsNo explicit noise predictionConversion is fragile near \(t=1\)
k-Diff\(kx_0-(1-k)\epsilon\)One learned scalar shared globallySame mixture in every directionOne global compromise
AsymFlow\(x_0-P\epsilon\)Binary rank-\(r\) PCA maskNoise in selected directionsDirect velocity in the selected subspace
MFlow\(x_0-M\epsilon\)Continuous gain per PCA directionLearned soft allocationLearned direction by direction

What happens in practice?

The main empirical story is not that one fixed target wins everywhere. It is that the preferred target changes with the domain, representation, objective, and model scale.

SettingPatch / tubeletJiT / \(x_0\)AsymFlowMFlow
ImageNet
B + LPIPS, FID ↓\(16\times16\)3.423.403.33
H + LPIPS, FID ↓\(16\times16\)1.821.791.81
UCF-101
B, FVD ↓\(8\times8\times4\)152126122
L, FVD ↓\(8\times8\times4\)1119897
Text-to-image
1.1B, GenEval ↑\(16\times16\)0.88390.88320.8843
1.1B, DPG ↑\(16\times16\)84.6584.6184.70
1.1B, PRISM ↑\(16\times16\)73.673.973.9

For UCF-101, the tubelet is listed as spatial × spatial × temporal.

MFlow is competitive or better across these settings without a rank sweep, but the gains over strong fixed parameterizations are sometimes modest. That nuance matters. The value of a learnable parameterization is not only a lower headline metric; it is also the ability to adapt the target when the data, loss, representation, or model capacity changes.

Epilogue

Prediction targets are often introduced as interchangeable algebraic views of the same diffusion process. That statement is true for an exact model. It is incomplete for a finite one.

Clean-data prediction is easier to compress but harder to convert near clean data. Velocity prediction is well-conditioned but carries full-dimensional noise. k-Diff learns a global compromise. AsymFlow allocates a binary velocity subspace. MFlow turns that subspace into a continuous, learnable curve.

So the answer to the title is not velocity or the clean image. It is often velocity in some directions, the clean image in others, and something in between when the model benefits from a softer trade-off.

A collage of sample images generated by the MFlow XXL/16 model, including macro photography, portraits, interiors, animals, and architectural scenes.
Figure 3. Sample images generated by MFlow XXL/16, trained on 512p images (13B-parameter mixture-of-experts, 2.9B active parameters per token).

Under a fixed velocity budget, where should it go?

Suppose a finite model can support only a limited amount of velocity-like prediction. In MFlow language, constrain the soft velocity budget

\[r_{\mathrm{eff}}=\sum_{i=1}^D m_i=B.\]

Where should those gains be placed? There is an important caveat before doing any algebra: the exact velocity-space flow-matching objective is rotationally invariant. With an infinitely expressive model and perfect optimization, it does not prefer one orthonormal basis over another. A preference for leading PCA directions appears only after we add the two ingredients that matter in practice: finite capacity and a budget on how much Gaussian noise the network is asked to predict.

Let \(\tilde x_0=U^\top x_0\) and \(\tilde\epsilon=U^\top\epsilon\). Write \(\lambda_i=\mathbb E[\tilde x_{0,i}^2]\) for the directional data energy, ordered as \(\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_D\), while isotropic noise has equal energy in every direction:

\[ \mathbb E[\tilde x_{0,i}^2]=\lambda_i, \qquad \mathbb E[\tilde\epsilon_i^2]=1. \]

In direction \(i\), MFlow predicts \(\tilde u_{G,i}=\tilde x_{0,i}-m_i\tilde\epsilon_i\). If \(e_{G,i}=\hat{\tilde u}_{G,i}-\tilde u_{G,i}\), exact recovery gives

\[ e_{u,i} =\hat{\tilde u}_i-\tilde u_i =\frac{e_{G,i}}{\Delta_i(t)}, \qquad \Delta_i(t)=(1-t)+tm_i. \]

Therefore the ordinary velocity-space flow-matching loss can be written as

\[ \mathcal L_{\mathrm{FM}} =\mathbb E\!\left[ w(t)\sum_i\bigl(\hat{\tilde u}_i-\tilde u_i\bigr)^2 \right] =\mathbb E\!\left[ w(t)\sum_i\frac{e_{G,i}^2}{\bigl((1-t)+tm_i\bigr)^2} \right]. \]

A nonzero \(m_i\) improves clean-endpoint conditioning: for \(m_i=0\), the conversion divides by \(1-t\); for \(m_i>0\), \(\Delta_i(t)\to m_i\) as \(t\to1\). But that improvement is not free. Increasing \(m_i\) injects more Gaussian noise into the direct regression target. Assuming signal and noise are independent,

\[ \mathbb E[\tilde u_{G,i}^2] =\lambda_i+m_i^2, \qquad \underbrace{\frac{m_i^2}{\lambda_i}}_{\text{relative stochastic burden}} \text{ measures how much noise is added per unit of data energy.} \]

This makes the top-PCA allocation a natural default. Under a hard budget \(m_i\in\{0,1\}\) with \(\sum_i m_i=N\), choosing the leading \(N\) directions does two things at once:

The same conclusion has a soft version. As a simple finite-capacity proxy, distribute a budget \(B\) while minimizing the total relative stochastic burden:

\[ \min_{0\le m_i\le1} \sum_i\frac{m_i^2}{\lambda_i} \quad\text{subject to}\quad \sum_i m_i=B. \]

Ignoring the box constraint for a moment, the Lagrangian condition gives

\[ \frac{2m_i}{\lambda_i}+\eta=0 \quad\Longrightarrow\quad m_i=\frac{B\lambda_i}{\sum_j\lambda_j}. \]

After clipping any saturated gains at one and reallocating the remainder, the solution still assigns at least as much gain to a higher-variance direction as to a lower-variance one. In words: spend scarce velocity capacity where it buys conditioning for the most data variation, and leave low-variance, predominantly off-manifold directions closer to clean-data prediction.

This is a prior, not a universal theorem. The standard flow loss itself does not say “use the first PCs.” The argument depends on an isotropic-noise model and a finite-capacity proxy. Model approximation error can be anisotropic, a non-leading PCA block can occasionally work better, and auxiliary objectives can value directions by visual function rather than by variance. That is precisely why MFlow initializes from variance but keeps every \(m_i\) learnable.

What do we mean by a velocity budget? We use \(B_{\mathrm{soft}}:=\operatorname{tr}(M)=\sum_i m_i\), a soft effective rank that equals the exact rank for binary AsymFlow masks but varies continuously for MFlow; this continuity, and the absence of an arbitrary threshold, is why it is the quantity constrained in the derivation above. For reporting, one could instead define \(B_{\mathrm{cutoff}}(\tau):=\sum_i\mathbf 1[m_i>\tau]\), which counts only directions whose gains exceed a tolerance such as \(\tau=10^{-2}\). Under the same convention, k-Diff is still all-or-nothing in rank: because its shared noise coefficient is \(1-k\), \(B_{\mathrm{cutoff}}^{\mathrm{k\text{-}Diff}}(\tau)=D\) when \(1-k>\tau\) and \(0\) otherwise (so at \(\tau=10^{-2}\), the switch occurs at \(k=0.99\)). Thus k-Diff can shrink the coefficient globally, but cannot allocate a fixed-rank budget to selected directions.

Why does LPIPS bend the M-curve?

On ImageNet, adding a perceptual loss improves all three parameterizations in the controlled B-level comparison, but it helps MFlow the most:

ParameterizationFID without LPIPS ↓FID with LPIPS ↓Change
JiT / \(x_0\)3.703.42−0.28
AsymFlow3.593.40−0.19
MFlow3.743.33−0.41

The training objective is now schematically

\[ \mathcal L =\mathcal L_{\mathrm{FM}} +\beta\,d_{\mathrm{LPIPS}}(\hat x_0,x_0), \qquad \hat{\tilde x}_{0,i} =\frac{(1-t)\hat{\tilde u}_{G,i}+m_i\tilde x_{t,i}}{(1-t)+tm_i}. \]

Unlike squared error in velocity space, LPIPS is a nonlinear, spatial feature-space distance. It is neither diagonal in the patch-PCA basis nor ordered by PCA variance.

So once LPIPS is added, there is no reason for the optimal gains to decrease monotonically with \(\lambda_i\). The appendix shows exactly such a departure: without LPIPS, the leading gains roughly follow the variance ordering; with LPIPS, PC0 and PC1 stay high, PC2 and PC3 form a pronounced dip, and PC4 rises again.

Top 64 modes of the uncentered 16 by 16 by 3 patch PCA basis, labeled PC0 through PC63 with eigenvalues.
Figure 3. Top 64 modes of the uncentered \(16\times16\times3\) patch-PCA basis. The LPIPS-related non-monotonicity aligns with the spatial role of the leading modes: PC0–PC1 are patch-global, PC2–PC3 are spatial gradients, and PC4 returns to a patch-global chromatic mode. Eigenvector signs are arbitrary. Click for full resolution.

The modes make the dip less mysterious. PC0 is dominated by low-frequency patch tone, and PC1 is approximately uniform across the patch while changing color. PC2 and PC3 instead encode vertical and horizontal spatial gradients. PC4 then returns to an approximately uniform chromatic mode. The structural sequence is therefore

\[ \text{patch-global},\quad \text{patch-global},\quad \text{spatial},\quad \text{spatial},\quad \text{patch-global}, \]

which mirrors the learned LPIPS gain pattern

\[ \text{high},\quad \text{high},\quad \text{low},\quad \text{low},\quad \text{high}. \]

The rebound at PC4 cannot be explained by variance alone: its directional variance is \(\lambda_4=3.74\), below both \(\lambda_2=9.41\) and \(\lambda_3=7.93\), yet LPIPS gives PC4 the larger gain. Nor is the behavior simply “LPIPS prefers color,” because several later chromatic modes are spatially non-uniform and do not receive the same increase.

What the experiment establishes: LPIPS causes the non-monotonic M-curve, and the high-gain rebound aligns with patch-global modes.

Open questions

Should the gains depend on time?

MFlow learns one gain per direction, shared across timesteps. In principle a time-dependent \(m_i(t)\) could allocate conditioning differently near the noise and clean endpoints. We tried it: we split \(M\) into time buckets of 50 steps (out of the 1000 flow-matching steps) and learned a separate gain vector per bucket. In our experiments this produced much worse FID, so we kept a single time-shared gain per direction.

Should the basis be learned?

PCA gives an interpretable, fixed coordinate system. A learned or data-dependent basis might discover directions better aligned with the model's approximation error, though it would make the method less simple.

Can gain learning be decoupled from network learning?

Jointly updating the network and \(M\) creates a moving direct target. Staged, bilevel, or slower-timescale gain optimization may preserve the adaptive benefit while reducing co-adaptation.


Acknowledgements

I would like to thank my co-authors, institutes, and funding sources for supporting this work.

The idea for this blog was inspired by the amazing blog To be (parametric) or not to be, that is the question. by Yue Zhao.

This blog is based on work: Rethinking the Design Space of Vanilla Pixel Diffusion Transformers (Mukhopadhyay et al., 2027).

References

  1. Li and He. Back to Basics: Let Denoising Generative Models Denoise. CVPR 2026. (JiT)
  2. Karras et al. Elucidating the Design Space of Diffusion-Based Generative Models. NeurIPS 2022. (EDM)
  3. Jin and Wang. Revisiting Diffusion Model Predictions Through Dimensionality. 2026. (k-Diff)
  4. Chen et al. Asymmetric Flow Models. 2026. (AsymFlow)
  5. Mukhopadhyay et al. Rethinking the Design Space of Vanilla Pixel Diffusion Transformers. 2027. (MFlow)