1 Abstract
In this section, we derive an expression for the space-variable gradient of the score function associated with the density of the stochastic interpolant, in terms of some covariance matrix.
2 Main Result
The following theorem is of main interest,
Theorem 1 Let \(s: [0,1] \times \mathbb R^{d} \rightarrow \mathbb R^{d}\) denote the score function associated with the density \(\rho_t\) of the stochastic interpolant \(X_t = a_t \, X_1 + b_t \, X_0\), where \((a, b)\) denotes an arbitrary FM schedule, \(X_0 \sim \mathcal N(0, I_d)\) and \(X_1 \sim \rho_1 = p_{\textrm{data}}\), independent \(d\)-dimensional random variables. For all \((t,x) \in (0,1)\times \mathbb R^{d}\), it holds that,
\[ \nabla_x s(t, x) = \dfrac{a_t^{2}}{b_t^{4}} \, \textrm{Cov}(X_1 \mid X_t = x) - \dfrac{1}{b_t^{2}} \, I_d \]
where \(\textrm{Cov}(X_1 \mid X_t = x)\) denotes the conditional expectation, \(\mathbb E_x\left[ \left(X_1 - \mathbb E_x[X_1]\right) \, \left( X_1 - \mathbb E_x[X_1] \right)^{T} \right]\)
Proof. Taking the gradient, w.r.t \(x\), on the expression of Lemma 1, yields,
\[ \nabla_x s(t, x) = -\nabla_x^{2} \log{} \rho_t(y \mid x) - \dfrac{1}{b_t^{2}} \, I_d \tag{1}\]
Substituting the result of Lemma 3 in Equation 1 gives exactly what we want to prove. \(\blacksquare\)
3 Lemmas
The lemmas that are used to prove the aforementioned theorem are presented in this section.
The first one is a basic expression that relates the score functions associated with the densities \(\rho_t(\cdot )\) and \(\rho_t(\, \cdot \mid x)\), for any \(x\in \mathbb R^{d}\).
Lemmas 2 and 3 provide a characterization of the aforementioned conditional density and an expression for the space-variable gradient of the score function associated with it.
Lemma 1 Assume the hypotheses of Theorem 1. For any \(y \in \mathbb R^{d}\), one has,
\[ s(t, x) = -\nabla_x \log{} \rho_t(y\mid x) - \dfrac{x - a_t y}{b_t^{2}}, \ \ \forall \, (t, x) \in (0,1) \times \mathbb R^{d} \]
Proof. For any \(t \in (0,1)\) and \(x,y \in \mathbb R^{d}\), Bayes theorem implies that
\[ \rho_t(x) = \dfrac{\rho_1(y) \, \rho_t (x \mid y)} {\rho_t(y\mid x)} \]
which in turn yields,
\[ \log{} \rho_t(x) = \log {} \rho_1(y) + \log {} \rho_t(x \mid y) - \log {} \rho_t (y\mid x) \]
Taking the gradient of this expression, w.r.t \(x\), one writes
\[ s(t, x) = -\nabla_x \log{} \rho_t(y \mid x) + \nabla_x \log{} \rho_t (x \mid y) \tag{2}\]
Now, the regular conditional distribution of \(X_t\) given the sigma-field \(\sigma(X_1)\) is gaussian with mean \(a_t y\) and covariance matrix \(b_t^{2} \, I_d\). Therefore,
\[ \begin{align*} \nabla_x \log {} \rho_t (x \mid y) & = \nabla_x \, \left[ -\dfrac{d}{2} \, \log {} (2\pi) - d \, \log{} (b_t) -\dfrac{||x - a_t y||^{2}} {2 \, b_t^{2}} \right] \\ & = -\dfrac{\cancel{2} \, (x - a_t y)}{\cancel{2} \, b_t^{2}} \end{align*} \]
Substituting the previous equation in Equation 2 gives the desired expression. \(\blacksquare\)
Lemma 2 For any \(t\in (0,1)\), the family of conditional densities \(\left\{\rho_t(\, \cdot \mid x) : x \in \mathbb R^{d}\right\}\) is exponential.
Proof. For any \(t \in (0,1)\) and \(x, y \in \mathbb R^{d}\) write,
\[ \begin{align*} \rho_t(y \mid x) & = \rho_1(y) \, \rho_t(x)^{-1} \, (2\pi)^{-d/2} \, b_t^{-d} \, \exp{} \left(-\dfrac{||x - a_t \, y||^{2}}{2b_t^{2}}\right) \\ & = \rho_1(y) \, \rho_t(x)^{-1} \, (2\pi)^{-d/2} \, b_t^{-d} \, \exp{} \left(-\dfrac{||x||^{2}}{2b_t^{2}} + \dfrac{a_t}{b_t^{2}} \left<x, y\right> -\dfrac{a_t^{2} \, ||y||^{2}}{2 \, b_t^{2}}\right) \end{align*} \tag{3}\]
Set \(\theta = \frac{a_t}{b_t^{2}} \, x\). Let \(h, A\) be functions from \(\mathbb R^{d}\) to \(\mathbb R\) s.t. \(h(y) = \rho_1(y) \, \exp{} \left(-\dfrac{a_t^{2} \, ||y||^{2}}{2b_t^{2}}\right)\) and \(A(\theta) = \log {} \, \int_{\mathbb R^{d}} h(y) \, \exp{} \left<y, \theta\right> \, dy\)
The aforementioned density is written, in terms of \(\theta\), as,
\[ \rho_t(y \mid \theta) = h(y) \, \exp{} \left(\left<y, \theta\right> - A(\theta)\right) \tag{4}\]
Indeed, for any \(x\in \mathbb R^{d}\), the following identity is true
\[ \begin{align*} \rho_t(x) & = \int_{\mathbb R^{d}} \, \rho_t(x \mid y) \, \rho_1(y) \, dy \\ & = (2\pi)^{-d/2} \, b_t^{-d} \, \exp{} \left(-\dfrac{||x||^{2}}{2b_t^{2}}\right) \, \int_{\mathbb R^{d}} \, \rho_1(y) \cdot \exp{} \left( \dfrac{a_t}{b_t^{2}} \, \left<x, y\right> - \dfrac{a_t^{2}}{2b_t^{2}} \, ||y||^{2} \right) \\ & = (2\pi)^{-d/2} \, b_t^{-d} \, \exp{} \left(-\dfrac{||x||^{2}}{2b_t^{2}}\right) \, \int_{\mathbb R^{d}} \, h(y) \, \exp{} \left<y, \theta\right> \, dy \end{align*} \]
and substituting this expression in Equation 3, we derive Equation 4.
Obviously, \(h\) is non-negative and the support of the density does not depend on \(\theta\). Also, the log-partition function is finite in all of \(\mathbb R^{d}\). Indeed for any \(y,\theta \in \mathbb R^{d}\) one has,
\[ h(y) \, \exp {} \left<y, \theta\right> = \rho_1(y) \, \exp{} \left(-\dfrac{a_t^{2}}{2b_t^{2}} \, || y - \frac{b_t^{2} \, \theta}{a_t^{2}} ||\right) \, \exp{} \left(\dfrac{b_t^{2} \, ||\theta||^{2}}{2 a_t^{2}}\right) \]
and integrating w.r.t \(y\) yields,
\[ e^{A(\theta)} \leq \exp{} \left(\dfrac{b_t^{2} \, ||\theta||^{2}}{2 \, a_t^{2}}\right) \ \implies A(\theta) \leq \dfrac{b_t^{2} \, ||\theta||^{2}}{2 \, a_t^{2}} < \infty \ \ \textrm{for any} \ \theta \in \mathbb R^{d} \]
which, in turn, implies the desired result. \(\ \ \ \blacksquare\)
Lemma 3 For any \(t\in (0,1)\) and \(x, y \in \mathbb R^{d}\), a covariance matrix expression of the log-Hessian matrix \(\nabla_x^{2} \log{} \rho_t(y \mid x)\) is given by
\[ \nabla_x^{2} \log{} \rho_t(y \mid x) = -\dfrac{a_t^{2}}{b_t^{4}} \, \textrm{Cov}\left(X_1 \mid X_t = x\right) \]
Proof. Due to Lemma 2, we can write the given density as in Equation 4. Taking the logarithm of this equation, we write,
\[ \log \, \rho_t(y \mid \theta) = \log{} h(y) + \left<y,\theta\right> - A(\theta) \tag{5}\]
Define \(Z: \mathbb R^{d} \rightarrow \mathbb R\) with \(Z(\theta) = \int_{\mathbb R^{d}} \, h(y) \, \exp{} \left<y, \theta\right> \, dy\).
As established in Lemma 2, \(\{\theta: Z(\theta) < \infty\} = \mathbb R^{d}\). (Theorem 2.4, W.Keener 2010) implies that \(Z\) is continuous and has continuous partial derivatives of all orders for \(\theta \in \mathbb R^{d}\). Furthermore, those derivatives can be computed by differentiation under the integral sign. Via the chain rule, the same thing can be said for the log-partition function.
Differentiating Equation 5 with respect to \(x\), we get,
\[ \nabla^{2}_x \log{} \, \rho_t(y \mid x) = - \nabla^{2}_x A(\theta) = -\dfrac{a_t^{2}}{b_t^{4}} \, \nabla^{2}_{\theta} \, A(\theta) \]
Invoking Theorem 1 from A property regarding exponential families gives the desired result. \(\ \ \ \blacksquare\)
Takeaway: The proofs above are similar in structure to (Lemma 18, Gao et al. 2024). To make things clearer we emphasized that the natural parameter space \(\Xi\), as defined in A property regarding exponential families, is all of \(\mathbb R^{d}\), which allowed us to invoke (Theorem 2.4, W.Keener 2010) and (Theorem 1, A property regarding exponential families) to finish off the proof of Lemma 3.