Exact Transport Laws under Joint Gradient Flow
Kevin L. Brown
August 2026
This paper develops an exact mathematical description of how task-relative representation quality evolves during learning, separating improvement in a model’s predictions from improvement in the information contained in its learned representation.
The starting point is a distinction that is often blurred in neural-network training. A predictor can improve because its current head is becoming better fitted, because its internal representation is becoming more useful for the task, or through some combination of both. These processes need not move in the same direction.
The paper studies this problem in a population teacher–student model with a learned linear representation and jointly trained linear head. For a representation matrix WW, the best achievable risk using that representation defines its representation-level Bayes excess relative to the unrestricted Bayes predictor. The scalar excess-risk concept itself is not claimed as new. The contribution is an exact matrix-valued lift of this quantity and its dynamics under simultaneous head–representation gradient flow.
The central object is the positive-semidefinite representation Bayes-excess matrix
S(W)=B(I−PW)B⊤,S(W)=B(I-P_W)B^\top,
where BB defines the teacher map and PWP_W projects onto the row space of the learned representation. Its trace measures the task-relevant information that the representation has not yet made recoverable by an optimal linear head.
On every constant-rank interval, the paper derives the exact transport law
S˙=−ηW(A⋆A⊤S+SAA⋆⊤),\dot S = -\eta_W \left( A^\star A^\top S + SAA^{\star\top} \right),
where A⋆A^\star is the instantaneous Bayes-optimal head supported by the current representation.
Several consequences follow.
First, the Bayes-excess matrix evolves by congruence on regular rank strata, preserving its rank at every finite regular time. Exact Bayes sufficiency is therefore invariant once reached, while an initially insufficient representation cannot cross smoothly into exact sufficiency at finite regular time.
Second, the scalar Bayes excess obeys an exact alignment law. This proves that population prediction loss can decrease while representation Bayes excess increases. In other words, a network can become a better predictor while simultaneously learning a worse task representation.
Third, the paper derives an exact cumulative-alignment criterion for asymptotic convergence. Sustained positive task-weighted alignment between the actual head and the instantaneous optimal head produces exponential decay of representation Bayes excess, while insufficient cumulative alignment leaves a nonzero asymptotic floor.
The analysis also compares joint training with Variable Projection, where the head is optimized out at every instant. Under this condition, representation Bayes excess is monotone nonincreasing, identifying head lag or misalignment as the mechanism permitting representation degradation during ordinary joint training.
Additional results address symmetric and zero-head initialization, full-rank preservation, covariance-weighted static geometry, discrete-step local correspondence, and isolated analytic rank-loss events. The underlying parameter trajectory remains a smooth polynomial gradient flow; the Bayes-excess matrix is treated as a derived algebraic observable that may exhibit isolated point spikes when representation rank changes.
The paper does not claim a nonlinear-network theorem, finite-sample generalization result, or empirical validation. Its contribution is a scoped exact theory showing that optimization of prediction error and optimization of representation sufficiency are mathematically distinct dynamical processes.
