1 Abstract
In this section, we will focus on proving two key inequalities regarding the variance of a \(C^{1}\) transformation of a random variable \(X\), defined on a probability space \((\Omega, \mathcal F, \mathbb P)\), where \(\mathbb P \ll \lambda_d\), and it’s associated density is \(p(x) = e^{-V(x)}, \ x \in \mathbb R^{d}\) for \(V \in C^{2}(\mathbb R^{d})\).
In the following statements we use \(\left<\cdot, \cdot\right>\) to denote the standard inner product in \(\mathbb R^{d}\). \(||\cdot||\) will denote the induced norm for vectors in \(\mathbb R^{d}\) and the induced operator norm for matrices in \(\mathbb R^{d\times d}\).
Additional hypotheses to be imposed on \(V\), are shown below,
\[ ||\nabla V||^{2}, ||\nabla^{2} V|| \in \mathcal{L}^{1}(\mathbb P) \ \ \textrm{and} \ \ \mathbb E_{\mathbb P} [\nabla^{2} V] \succ 0 \tag{1}\]
2 Main Results
2.1 Cramer-Rao Inequality
Theorem 1 Let \(X: \Omega \rightarrow \mathbb R^{d}\) be a random variable defined on \((\Omega, \mathcal F, \mathbb P)\) and let \(h \in C^{1}(\mathbb R^{d})\) be a real-valued function, such that \(h^{2}, \nabla h \in \mathcal{L}^{1}(\mathbb P)\). Also, assume the statements in 1, regarding the \(C^{2}\) potential function \(V\). One has,
\[ \textrm{Var}_{\mathbb P}(h(X)) \geq \left<\mathbb E_{\mathbb P} \, \nabla h, \, \left(\mathbb E_{\mathbb P}[\nabla^{2} V]\right)^{-1} \, \mathbb E_{\mathbb P} \, \nabla h\right> \tag{2}\]
We will first prove the result for continuously differentiable functions \(h\) that have compact support, i.e. for \(h\in C_{c}^{1}(\mathbb R^{d})\). Note that in this setting, integrability assumptions on \(h\) and it’s gradient are not necessary.
Lemma 1 Let \(X: \Omega \rightarrow \mathbb R^{d}\) be a random variable defined on \((\Omega, \mathcal F, \mathbb P)\) and let \(h \in C^{1}_{c}(\mathbb R^{d})\) be a real-valued function. Also, assume the statements in 1, regarding the \(C^{2}\) potential function \(V\). Then, 2 holds.
Proof. We start off by showing the following equalities,
\[ \mathbb E_{\mathbb P} \nabla V = 0 \ \ \textrm{and} \ \ \mathbb E_{\mathbb P} \nabla h = \int \, h \, \nabla V \, d\mathbb P \]
Equivalently, it suffices to prove that,
\[ \mathbb E_{\mathbb P} \, \partial_i \, V = 0 \ \ \textrm{and} \ \ \mathbb E_{\mathbb P} \, \partial_i \, h = \mathbb E_{\mathbb P} [ h \, \partial_i \, V] \ \ \textrm{for any} \ i\in \{1,\dots,d\} \tag{3}\]
Since,
\[ \begin{align*} \begin{split} \mathbb E_{\mathbb P} \, |\partial_i V| \leq \left(\mathbb E_{\mathbb P}[|\partial_i \, V|^{2}]\right)^{1/2} \leq C^{1/2} \, \left(\mathbb E_{\mathbb P}||\nabla V||^{2}\right)^{1/2} < \infty \end{split} \end{align*} \]
for some positive constant \(C\), the function \(x \rightarrow \partial_i V(x) \, e^{-V(x)}\) is in \(\mathcal L^{1}(\mathbb R^{d})\), and Fubini’s theorem, implies,
\[ \infty > \mathbb E_{\mathbb P} \, \partial_i \, V = \int_{\mathbb R^{d-1}} \, \left[ \int_{\mathbb R} \, \partial_i V(x_i, x_{-i}) e^{-V(x_i, x_{-i})} \, dx_i \right] \, dx_{-i} \]
where \(x_{-i}\) denotes the vector \(x\) without the \(i\)-th coordinate. Due to the finiteness of the previous integral, we have,
\[ \int_{\mathbb R} \partial_i V(x_i, x_{-i}) e^{-V(x_i, x_{-i})} \, dx_i < \infty \ \ \textrm{for a.e} \ x_{-i} \in \mathbb R^{d-1} \]
Via the same argument, we can also deduce that,
\[ \int_{\mathbb R} e^{-V(x_i, x_{-i})} \, dx_i < \infty \ \ \textrm{for a.e} \ x_{-i} \in \mathbb R^{d-1} \]
We can see that, for a.e \(x_{-i}\), the function \(x_i \rightarrow g(x_i) := e^{-V(x_i, x_{-i})}\), is continuously differentiable and integrable along with it’s derivative. The auxiliary result in Proposition 2 shows that for such functions of one variable, the limits at positive and negative infinity exist, and are equal to 0. Therefore, for a.e \(x_{-i} \in \mathbb R^{d-1}\), one has,
\[ \begin{align*} \int_{\mathbb R} \, \partial_i V(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})} \, dx_i & = -\lim_{R\rightarrow \infty} \, \int_{-R}^{R} \, \partial_i \left(e^{-V(x_i, x_{-i})}\right) \, dx_i \\ & = - \lim_{R \rightarrow \infty} g(R) + \lim_{R \rightarrow \infty} \, g(-R) = 0 \end{align*} \]
This, in turn, yields,
\[ \mathbb E_{\mathbb P} \, \partial_i V = \int_{\mathbb R^{d-1}} \, 0 \, dx_{-i} = 0 \]
To get the second equality in 3, one can see that it suffices to prove,
\[ \int_{\mathbb R^{d}} \, \partial_i \left(h(x) \, e^{-V(x)}\right) \, dx = 0 \ \ \textrm{for any} \ i \in \{1,\dots, d\} \]
Define \(\phi: \mathbb R^{d} \rightarrow \mathbb R\) with \(\phi(x) = h(x) \, e^{-V(x)}\), \(x \in \mathbb R^{d}\). The function \(\partial_i \, \phi\) is continuous and compactly supported, which means that it is integrable with respect to \(\lambda_d\). Fubini’s theorem implies,
\[ \int_{\mathbb R^{d}} \, \partial_i \left(h(x) \, e^{-V(x)}\right) \, dx = \int_{\mathbb R^{d-1}} \, \left[\int_{\mathbb R} \, \partial_i \left(h(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})}\right) \, dx_i\right] \, dx_{-i} \]
For every \(x_{-i} \in \mathbb R^{d-1}\), the function \(\phi\), when viewed exclusively as a function of \(x_i\), is continuously differentiable with compact support. The auxiliary result in Proposition 3 implies that,
\[ \int_{\mathbb R} \, \partial_i \left(h(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})}\right) \, dx_i = 0 \ \ \textrm{ for any} \ x_{-i} \in \mathbb R^{d-1} \]
which, in turn, yields,
\[ \int_{\mathbb R^{d}} \, \partial_i \left(h(x) \, e^{-V(x)}\right) \, dx = \int_{\mathbb R^{d-1}} \, 0 \, d x_{-i} = 0 \]
Now, to recover the variance of \(h \circ X\), under \(\mathbb P\), we write,
\[ \mathbb E_{\mathbb P} \, \nabla h = \int \, h \nabla V \, d\mathbb P = \int \, \left(h - \mathbb E_{\mathbb P} \, h\right) \, \nabla V \, d\mathbb P \]
The matrix \(A := \mathbb E_{\mathbb P} [\nabla^{2} \, V] \in \mathbb R^{d\times d}\) is a symmetric and invertible matrix with finite entries, due to the statements in 1. Multiplying the previous equation by \(A^{-1}\) and taking the inner product with \(\mathbb E_{\mathbb P} \, \nabla h\), one has,
\[ \begin{align*} \left<\mathbb E_{\mathbb P} \, \nabla h \, , A^{-1} \, \mathbb E_{\mathbb P} \, \nabla h\right> & = \left<\mathbb E_{\mathbb P} \, \nabla h \, , A^{-1} \, \int \left(h - \mathbb E_{\mathbb P} \, h\right) \, \nabla V \, d\mathbb P\right> \\ & = \int \, \left(h - \mathbb E_{\mathbb P} \, h\right) \, \left< \nabla V \, , A^{-1} \, \mathbb E_{\mathbb P} \, \nabla h\right> \, d\mathbb P \\ & := \int \, \left(h - \mathbb E_{\mathbb P} \, h\right) \left<\nabla V \, , c\right> \, d\mathbb P \end{align*} \]
The Cauchy-Schwarz inequality implies that
\[ \left<\mathbb E_{\mathbb P} \, \nabla h \, , A^{-1} \, \mathbb E_{\mathbb P} \, \nabla h\right> \leq (\textrm{Var}_{\mathbb P} \, h)^{1/2} \, \left[\int \, \left< \nabla V \, , c\right>^{2} \, d\mathbb P\right]^{1/2} \tag{4}\]
The validity of the previous inequality is due to the fact that \(h, h^{2} \in \mathcal L^{1}(\mathbb P)\), since \(h\in C^{1}_{c}(\mathbb R^{d})\), and due to
\[ \left<\nabla V, c\right>^{2} \leq ||\nabla V||^{2} \, ||c||^{2} \]
which gives \(\left< \nabla V \, , c \right>^{2} \in \mathcal{L}^{1}(\mathbb P)\). Also, notice that,
\[ \int \, \left< \nabla V \, , c\right>^{2} \, d\mathbb P = \left< \int \, \nabla V \, (\nabla V)^{T} \, d \mathbb P \, c \, , c\right> = \left<Ac \, , c\right> = \left<\mathbb E_{\mathbb P} \, \nabla h \, , A^{-1} \, \mathbb E_{\mathbb P} \, \nabla h\right> \]
due to the identity presented in Proposition 1.
Finally, assume that the quantity on the RHS of 2 is strictly positive. Otherwise, the result we want holds trivially. Dividing 4 by the square root of that quantity, and squaring both sides of the expression we obtain yields the desired inequality. \(\, \, \, \blacksquare\)
We now proceed with the proof of the theorem.
Proof. Theorem 1
Consider the smooth cutoff function \(\chi \in C^{\infty}_{c}(\mathbb R^{d})\) s.t.
\[ \chi(x) = 1 \ \ \textrm{inside} \ \overline{B}_1 \ , \ \chi(x) = 0 \ \ \textrm{outside} \ B_2 \ , \ 0 \leq \chi \leq 1 \ , \ ||\nabla \chi||_{\infty} \leq C_0 \] for some positive constant \(C_0\). Such a function exists via mollification of an indicator function. For more details we refer the reader to a standard functional analysis textbook.
Consider the sequence of functions \((\chi_n)_{n \geq 1}\), with,
\[ \chi_n(x) = \chi(x/n) \ , \ x\in \mathbb R^{d} \]
Fix \(n \in \mathbb N\). The function \(\chi_n\) satisfies the following properties,
\[ \chi_n(x) = 1 \ \ \textrm{inside} \ \overline{B}_n \ , \ \chi_n(x) = 0 \ \ \textrm{outside} \ B_{2n} \ , \ 0 \leq \chi_n \leq 1 \ , \ ||\nabla \chi_n||_{\infty} \leq \frac{C_0}{n} \]
Obviously, \(\chi_n\) is smooth, and
\[ \{\chi_n \neq 0\} \subseteq B_{2n} \ \implies \textrm {supp} \, \chi_n \subseteq \overline{B}_{2n} \]
implies that \(\chi_n\) has compact support, by the Heine-Borel theorem. Therefore, \(\chi_n \in C^{\infty}_{c} (\mathbb R^{d})\). Now, define the sequence of functions \((h_n)_{n \geq 1}\) with,
\[ h_n(x) = \chi_n(x) \, h(x) \ , \ x\in \mathbb R^{d} \]
Fix \(n\in \mathbb N\). The function \(h_n\) is continuously differentiable since \(\chi_n\) and \(h\) are both continuously differentiable. Also,
\[ \{h_n \neq 0\} \subseteq \{\chi_n \neq 0\} \subseteq B_{2n} \ \implies \textrm{supp} \, h_n \subseteq \overline{B}_{2n} \]
implies that \(h_n\) has compact support. Therefore, \(h_n \in C^{1}_{c}(\mathbb R^{d})\). Via the result of Lemma 1 one has,
\[ \textrm{Var}_{\mathbb P} \, h_n \leq \left<\mathbb E_{\mathbb P} \nabla h_n \, , \, \left(\mathbb E_{\mathbb P}[\nabla^{2} \, V]\right)^{-1} \, \mathbb E_{\mathbb P} \nabla h_n\right> \ , \ \ \textrm{for any} \ n\in \mathbb N \]
It suffices to justify why we can pass limits on the left and right hand side of the previous inequality.
We have that \(h_n\) and \(h_n^{2}\) converge pointwise to \(h\) and \(h^{2}\) respectively, and \(|h_n|, |h_n|^{2} \leq |h|, |h|^{2} \in \mathcal{L}^{1}(\mathbb P)\). The dominated convergence theorem implies
\[ \mathbb E_{\mathbb P} \, h_n \rightarrow \mathbb E_{\mathbb P} \, h \ \ \textrm{and} \ \ \mathbb E_{\mathbb P} \, h_n^{2} \rightarrow \mathbb E_{\mathbb P} \, h^{2} \]
which leaves us with, \(\textrm{Var}_{\mathbb P} \, h_n \rightarrow \textrm{Var}_{\mathbb P} \, h\). For any \(n \in \mathbb N\), the product rule yields
\[ \nabla h_n = \nabla \chi_n \, h + \chi_n \, \nabla h \]
and
\[ |\mathbb E_{\mathbb P}[\nabla \chi_n \, h]| \leq \int \, |\nabla \chi_n| \, |h| \, d\mathbb P \leq \int \, ||\nabla \chi_n||_{\infty} \, |h| \, d\mathbb P \leq \dfrac{C_0 \, \mathbb E_{\mathbb P} \, |h|}{n} \rightarrow 0 \]
since, \(\mathbb E_{\mathbb P} \, |h| < \infty\). Now, the sequence of functions \(\left(\chi_n \, \nabla h\right)_{n\geq 1}\) converges pointwise to \(\nabla h\) and \(|\chi_n \, \nabla h| \leq |\nabla h| \in \mathcal{L}^{1}(\mathbb P)\). The dominated convergence theorem implies
\[ \mathbb E_{\mathbb P}[\chi_n \, \nabla h] \rightarrow \mathbb E_{\mathbb P} \, \nabla h \]
So, \(\mathbb E_{\mathbb P} \, \nabla h_n \rightarrow \mathbb E_{\mathbb P} \, \nabla h\) and the fact that the function \(v \rightarrow \left<v, Av\right>\) is continuous in \(v\in \mathbb R^{d}\), where \(A\in \mathbb R^{d\times d}\) is a constant matrix, implies
\[ \left<\mathbb E_{\mathbb P} \nabla h_n \, , \, \left(\mathbb E_{\mathbb P}[\nabla^{2} \, V]\right)^{-1} \, \mathbb E_{\mathbb P} \nabla h_n\right> \rightarrow \left<\mathbb E_{\mathbb P} \nabla h \, , \, \left(\mathbb E_{\mathbb P}[\nabla^{2} \, V]\right)^{-1} \, \mathbb E_{\mathbb P} \nabla h\right> \]
and the result of the theorem has been proven. \(\ \ \ \blacksquare\)
2.2 Brascamp-Lieb Inequality
Let \(X: \Omega \rightarrow \mathbb R^{d}\) be a random variable defined on \((\Omega, \mathcal{F}, \mathbb P)\) and let \(h \in C_{c}^{\infty}(\mathbb R^{d})\) be a real-valued function. Assume that the potential function \(V\) belongs in \(C^{\infty}(\mathbb R^{d})\) and is strictly convex. One has,
\[ \textrm{Var}_{\mathbb P} \left(h(X)\right) \leq \mathbb E_{\mathbb P} \left[ \left<\nabla h \, , \, (\nabla^{2} V)^{-1} \, \nabla h\right>\right] \]
Proof. Omitted
3 Auxiliary Results
Proposition 1 Let \(\mathbb P\) be the probability measure defined in Section 1 and assume that the statements in 1 regarding the \(C^{2}\) potential function \(V\), hold. Then, one has,
\[ \int \, \nabla V \, (\nabla V)^{T} \, d\mathbb P = \int \, \nabla^{2} \, V \, d\mathbb P \]
Proof. For any \(i, j \in \{1,\dots, d\}\) it suffices to show that,
\[ \int \, \partial_i \, V \, \partial_j \, V \, d\mathbb P = \int \, \partial_{ij} \, V \, d\mathbb P \]
which is equivalent to,
\[ \int_{\mathbb R^{d}} \, \partial_i \left(\partial_j V(x) \, e^{-V(x)}\right) \, dx = 0 \tag{5}\]
Therefore, we will show that the integral in Equation 5 is finite and is equal to zero.
One writes,
\[ \begin{align*} \left|\int_{\mathbb R^{d}} \, \partial_i \left(\partial_j V(x) \, e^{-V(x)}\right) \, dx \right| & \leq \mathbb E_{\mathbb P}[|\partial_{ij} \, V|] + \mathbb E_{\mathbb P} [|\partial_i \, V \, \partial_j \, V|] \\ & \leq C \, \mathbb E_{\mathbb P}[||\nabla^{2} \, V||] + \mathbb E_{\mathbb P}[||\nabla \, V||^{2}] < +\infty \end{align*} \]
for some positive constant \(C\), since,
\[ \int \, |\partial_{ij} \, V| \, d\mathbb P \leq \int \, ||\nabla^{2} \, V||_{\textrm{max}} \, d\mathbb P \leq C \, \int \, ||\nabla^{2} \, V|| \, d\mathbb P = C \, \mathbb E_{\mathbb P}[||\nabla^{2} \, V||] \]
and
\[ \int \, |\partial_i \, V \, \partial_j \, V| \, d\mathbb P \leq \, \dfrac{1}{2} \, \int \, |\partial_i \, V|^{2} + |\partial_j \, V|^{2} \, d\mathbb P \leq \mathbb E_{\mathbb P}[||\nabla \, V||^{2}] \]
Now, Fubini’s theorem allows us to write,
\[ \infty > \int_{\mathbb R^{d}} \, \partial_i \left(\partial_j V(x) \, e^{-V(x)}\right) \, dx = \int_{\mathbb R^{d-1}} \, \left[\int_{\mathbb R} \, \partial_i\left(\partial_j \, V(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})}\right) dx_i\right] \, dx_{-i} \]
For a.e \(x_{-i} \in \mathbb R^{d-1}\) the function \(x_i \rightarrow w(x_i) := \partial_j \, V(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})}\) is continuously differentiable and integrable along with it’s derivative. Leveraging the result in Proposition 2 one calculates,
\[ \begin{align*} \int_{\mathbb R} \, \partial_i\left(\partial_j \, V(x_i, x_{-i}) \, e^{-V(x_i, x_{-i})}\right) dx_i & = \lim_{R \rightarrow \infty} \, \int_{-R}^{R} \, \partial_i \, w(x_i) \, dx_i \\ & = \lim_{R\rightarrow \infty} \, w(R) - \lim_{R \rightarrow \infty} \, w(-R) = 0 \end{align*} \]
and therefore,
\[ \int_{\mathbb R^{d}} \, \partial_i \left(\partial_j V(x) \, e^{-V(x)}\right) \, dx = \int_{\mathbb R^{d-1}} \, \left[\int_{\mathbb R} \, \partial_i \, w(x_i) \, dx_i\right] \, dx_{-i} = \int_{\mathbb R^{d-1}} \, 0 \, dx_{-i} = 0 \]
which yields the desired result. \(\blacksquare\)
Proposition 2 Consider a real-valued function \(h\in C^{1}(\mathbb R)\) s.t. \(h, h'\in \mathcal{L}^{1}(\mathbb R)\). The limits \(\lim_{x \rightarrow \pm \infty} \, h(x)\) exist and are equal to zero.
Proof. We will show that the limit of \(h\) as \(x\) tends to positive infinity exists and is equal to zero. The same result can be shown for the limit at negative infinity via similar arguments as the ones that follow.
The fundamental theorem of calculus allows us to write,
\[ h(x) = h(0) + \int_{0}^{x} \, h'(t) \, dt := h(0) + F(x) \ \ \textrm{for any} \ x > 0 \]
Let \(\varepsilon > 0\). Since \(h' \in \mathcal{L}^{1}(\mathbb R)\), there exists an \(M > 0\), s.t.
\[ \int_{M}^{\infty} \, |h'(t)| \, dt < \varepsilon \]
Therefore, for any \(M < x < y\), one has,
\[ |F(y) - F(x)| \leq \int_{x}^{y} \, |h'(t)| \, dt \leq \int_{M}^{\infty} \, |h'(t)| \, dt < \varepsilon \]
This implies the existence of \(\lim_{x\rightarrow \infty} \, F(x)\) which in turn implies the existence of \(\lim_{x\rightarrow \infty} \, h(x)\). Now suppose for the sake of contradiction that this limit is non-zero and equal to some real number \(L\). There exists an \(A > 0\) s.t.
\[ |h(x) - L| < \dfrac{|L|}{2} \implies |h(x)| > \dfrac{|L|}{2} \ \ \textrm{for any} \ x \geq A \]
and we arrive at the statement
\[ \int_{A}^{\infty} \, h(x) \, dx > \int_{A}^{\infty} \, \dfrac{|L|}{2} \, dx = \infty \]
which is contradictory to \(h \in \mathcal{L}^{1}(\mathbb R)\). \(\blacksquare\)
Proposition 3 Consider a real-valued function \(h \in C^{1}_{c}(\mathbb R)\). Then, \(h' \in \mathcal{L}^{1}(\mathbb R)\) and \(\int_{\mathbb R} \, h'(x) \, dx = 0\).
Proof. Since \(h'\) is a continuous function and it’s support is included in \(\textrm{supp} \, h\), we have \(h' \in C_{c}(\mathbb R)\). Therefore, \(h' \in \mathcal{L}^{1}(\mathbb R)\). Now, let \(K := \textrm{supp} \, h\). Due to the compactness of \(K\), there exists an \(R > 0\) s.t. \(K \subseteq (-R, R)\), which implies, \(h(R) = h(-R) = 0\). Also, \(h'\) vanishes outside of \([-R, R]\), and we get
\[ \int_{\mathbb R} \, h'(x) \, dx = \int_{[-R, R]} h'(x) \, dx + \int_{\mathbb R \setminus [-R, R]} \, h'(x) \, dx = h(R) - h(-R) + 0 = 0 \]
\(\blacksquare\)