1 Abstract
This post is written to help readers understand a key property regarding distributions whose densities form an exponential family parametrized on some \(\theta \in \mathcal{\Theta} \subseteq \mathbb R^{k}\).
2 Definitions
The following definitions are taken, with some minor notational changes, from (Section 2.1, W.Keener 2010).
Let \(\mu\) be a measure on \(\mathbb R^{d}\), let \(h: \mathbb R^{d} \rightarrow \mathbb R\) be a nonnegative function and let \(T: \mathbb R^{d} \rightarrow \mathbb R^{k}\) be a measurable function. For \(\theta \in \mathbb R^{k}\), define,
\[ A(\theta) = \log{} \int \, h(x) \, \exp{} \left<T(x), \theta\right> \, d\mu(x) \]
Whenever, \(A(\theta) < \infty\), the function \(p(\cdot | \theta)\) given by
\[ p(x \mid \theta) = h(x) \, \exp{} \left[\left<T(x), \theta\right> - A(\theta)\right], \ \ x\in \mathbb R^{d} \]
satisfies, \(\int p \, d\mu = 1\). This construction gives a family of probability densities indexed by \(\theta\).
The set \(\Xi = \{\theta: A(\theta) < \infty\}\) is called the natural parameter space, and the family of densities \(\{p(\cdot\mid \theta) : \theta \in \Xi\}\) is called a \(d\)-parameter exponential family in canonical form.
3 Main Result
The following result relates the expectation and covariance matrix of \(T\), under \(\mu\), to the derivatives of \(A\), as in Section 2.
Theorem 1 Let \((\Omega, \mathcal{F}, \mathbb P)\) be a probability space and consider a a random variable \(X\) taking values in the measure space \((\mathcal{X}, \mathcal{A}, \mu)\). Let \(p(\cdot \mid \theta)\) be the density of \(X\) with respect to \(\mu\). One has,
\[ \mathbb E_{\mu}[T(X)] = \nabla_{\theta} \, A(\theta) \ \ \textrm{and} \ \ \textrm{Cov}_{\mu} \, (T(X)) = \nabla^{2}_{\theta} \, A(\theta), \ \ \textrm{for all} \ \theta \in \Xi^{\circ} \]
Proof.
(Theorem 2.4, W.Keener 2010) implies that the function
\[ g(\theta) := \int \, h(x) \, \exp{} \left< T(x), \theta\right> \, d\mu(x) = e^{A(\theta)} \tag{1}\]
is continuous and has continuous partial derivatives of all orders for \(\theta \in \Xi^{\circ}\). Furthermore, these derivatives can be computed by differentiation inder the integral sign.
Differentiating Equation 1 with respect to some \(\theta_j\) yields,
\[ \dfrac{\partial A(\theta)}{\partial \theta_j} \, \cancel{e^{A(\theta)}} = \int \, h(x) \, T_j(x) \, \exp \left< T(x), \theta\right> \, d\mu(x) = \mathbb E_{\mu} [T_j(X)] \, \cancel{e^{A(\theta)}} \tag{2}\]
for all \(\theta\) in the interior of \(\Xi\).
Now, differentiating \(\frac{\partial g(\theta)}{\partial \theta_j}\) with respect to some \(\theta_k\), yields,
\[ \dfrac{\partial^{2} g(\theta)}{\partial \theta_k\, \partial \theta_j} = \dfrac{\partial^{2} A(\theta)}{\partial \theta_k\, \partial \theta_j} \, e^{A(\theta)} + \dfrac{\partial A(\theta)}{\partial \theta_j} \, \dfrac{\partial A(\theta)}{\partial \theta_k} \, e^{A(\theta)} \tag{3}\]
But,
\[ \dfrac{\partial^{2} g(\theta)}{\partial \theta_k \, \partial \theta_j} = \int \, h(x) T_j(x) \, T_k(x) \, \exp{} \left<T(x), \theta\right> \, d\mu(x) = \mathbb E_{\mu}[T_j(X) \, T_k(X)] \, e^{A(\theta)} \tag{4}\]
Substituting Equation 4 in Equation 3 yields,
\[ \dfrac{\partial^{2} A(\theta)}{\partial \theta_k \, \partial \theta_j} = \mathbb E_{\mu}[T_j(X) \, T_k(X)] - \mathbb E_{\mu} [T_j(X)] \, \mathbb E_{\mu} [T_k(X)] = \textrm{Cov}_{\mu}(T_j(X), T_k(X)) \tag{5}\]
for all \(\theta\) in the interior of \(\Xi\).
Equations 2 and 5 give the desired result. \(\ \ \ \blacksquare\)
4 References
For proofs utilizing Theorem 1 we refer the reader to A covariance equality for the gradient of the score function
The references that are not included in the note collection, are given below.