3.7 Probability distributions of variables and parameters
As we did in the previous chapter, we will keep the same assumption about the normality distribution of errors to be able to derive all the required distributions.
3.7.1 Probability distribution in terms of the unknown \(\sigma^2\)
All the distributions of our model variables and parameters are based on the assumption of normality of \(\varepsilon\).
- Distribution of \(y\):
Using the DGP model (3.2), we can derive the distribution of \(y\) from that of \(\varepsilon\) as follows:
\[\begin{align} \varepsilon\sim N(0,\sigma^2I)&\implies\notag\\ (y-X\beta)\sim N(0,\sigma^2I)&\implies\notag\\ y \sim N(X\beta,\sigma^2I) \tag{3.37} \end{align}\]
Which is also normally distributed with mean \(E(y)=X\beta\) and variance \(V(y)=\sigma^2I\).
- Distribution of \(\widehat\beta\):
Since the vector \(\widehat\beta\) is a function of the normally distributed variable \(\varepsilon\) (see (3.13)):
\[\begin{equation*} \widehat\beta=\beta+\big(X^tX\big)^{-1}X^t \varepsilon \end{equation*}\]
With mean \(E(\widehat\beta)=\beta\) (see 3.2.2) and variance \(Var(\widehat\beta)=\sigma^2\big(X^tX\big)^{-1}\) (see 3.2.3), then \(\widehat\beta\) is normally distributed:
\[\begin{equation} \widehat\beta\sim N(\beta,\sigma^2\big(X^tX\big)^{-1}) \tag{3.38} \end{equation}\]
- Distribution of \(\widehat y\):
Since \(\widehat y\) is a function of \(\widehat \beta\) in (3.8), It is normally distributed with mean:
\[\begin{equation*} E(\widehat y)=X\beta \end{equation*}\]
and variance:
\[\begin{align} Var(\widehat y)&=E\bigg[\bigg(\widehat y-E(\widehat y)\bigg)\bigg(\widehat y-E(\widehat y)\bigg)^t\bigg]\notag\\ &=E\bigg[\bigg(X\widehat\beta-X\beta\bigg)\bigg(X\widehat\beta-X\beta\bigg)^t\bigg]\notag\\ &=XE\bigg[\bigg(\widehat\beta-\beta\bigg)\bigg(\widehat\beta-\beta\bigg)^t\bigg]X^t\notag\\ &=XVar(\widehat\beta)X^t\notag\\ &=\sigma^2X\big(X^tX\big)^{-1}X^t \tag{3.39} \end{align}\]
That is:
\[\begin{equation} \widehat y\sim N\bigg(X\beta,\sigma^2X\big(X^tX\big)^{-1}X^t\bigg) \tag{3.40} \end{equation}\]
- Distribution of \(e\):
using \(e=y-\widehat y\), The mean of this variable is zero:
\[\begin{equation} E(e)=E(y)-E(\widehat y)=X\beta-X\beta=0 \end{equation}\]
And the variance is:
\[\begin{align*} Var(e)&=Var(y)+Var(\widehat y)\\ &=\sigma^2+\sigma^2X\big(X^tX\big)^{-1}X^t\\ &=\sigma^2\bigg(I_n+X\big(X^tX\big)^{-1}X^t\bigg) \end{align*}\]
That is:
\[\begin{equation} e\sim N\bigg(0,\sigma^2\big(I_n+X\big(X^tX\big)^{-1}X^t\big)\bigg) \end{equation}\]
- Distribution of \(s^2\):
We start driving this distribution from the expression (3.18), which has the symmetric and idempotent matrix \(M\) (see 3.3). Since the \(M\) matrix has the rank \(n-k\), its eigenvalues are either equal to one or to zero as follows:
\[\begin{equation} Mx=\lambda x\implies MMx=M(\lambda x)=\lambda(\lambda x)=\lambda^2x \end{equation}\]
Using the idempotent property of \(M\) (\(MM=M\)), we obtain:
\[\begin{align*} MMx=Mx&\implies \lambda^2x=\lambda x\\ &\implies (\lambda^2-\lambda)x=0\\ &\implies \lambda^2-\lambda=0\\ &\implies \lambda=0\quad ou \quad \lambda=1 \end{align*}\]
To figure out how much ones we have, we compute the trace of \(M\):
\[\begin{align*} Tr(M)&=Tr\bigg(I_n-X\big(X^tX\big)^{-1}X^t\bigg)\\ &=Tr(I_n)-Tr\bigg(\underbrace{X^tX\big(X^tX\big)^{-1}}_{I_k}\bigg)\\ &=n-k \end{align*}\]
Therefore, since the trace is equal to \(n-k\) then we have \(n-k\) of 1’s and \(k\) of 0’s, and hence the matrix has rank \(n-k\).
By transforming the expression (3.18) in standardized form by dividing the two sides by \(\sigma^2\) as follows:
\[\begin{equation} \frac{e^te}{\sigma^2}=\bigg(\frac{\varepsilon}{\sigma}\bigg)^tM\bigg(\frac{\varepsilon}{\sigma}\bigg) \end{equation}\]
Notice that the transformed random variable \(\frac{\varepsilon}{\sigma}\) follows the standard normal distribution with mean zero and variance \(I_n\). We know from probability theory (see 2.11) that the sum of squared standard normal variables is chi-square variable with degrees of freedom equals to the number of these variables. For our case, we have \(n-k\) (which is the rank of \(M\)) random squared standard variables, so it follows that:
\[\begin{equation} \frac{e^te}{\sigma^2}=(n-k)\frac{s^2}{\sigma^2}\sim \chi^2_{n-k} \tag{3.41} \end{equation}\]
3.7.2 Probability distribution in terms of the estimated variance \(s^2\)
All the previous distributions are based on \(\sigma^2\), which is usually unknown, and hence the different variances can not be computed. Using, however, the estimated \(s^2\) instead may yield in other types of distributions.
- Distribution of \(\widehat\beta\):
Let \(\alpha_{ij}\) be the elements of the matrix \((X^tX)^{-1}\). We have shown that each diagonal element of \(V(\widehat\beta)=\sigma^2(X^tX)^{-1}\) is the variance of a single estimator. For instance, the variance of \(\widehat\beta_i\) is \(\sigma^2\alpha_{ii}\).
Using 2.11 and (3.41), the distribution of each single estimator can be determined as follows:
\[\begin{equation} \frac{N(0,1)}{\sqrt{\chi^2_{n-k}\bigg/(n-k)}}=\frac{\widehat\beta_i-\beta_i\bigg/\sigma^2_{\widehat\beta_i}}{\sqrt{(n-k)s^2_{\widehat\beta_i}\bigg/\sigma^2_{\widehat\beta_i}(n-k)}}=\frac{\widehat\beta_i-\beta_i}{s_{\widehat\beta_i}}\sim t(n-k) \tag{3.42} \end{equation}\]
That means that \(\widehat\beta_i\) follows the student distribution with \(n-k\) degrees of freedom.
- Distribution of \(\widehat y\):
Let \(\gamma_{ij}\) be the elements of the hat matrix \(H\) so that the variance of the single variable \(\widehat y_i\) is equal to \(V(\widehat y_i)=\sigma^2\gamma_{ii}\) (where \(\gamma_{ii}\) are the diagonal elements of \(H\)). The distribution of \(\widehat y_i\) thus will be:
\[\begin{equation} \frac{N(0,1)}{\sqrt{\chi^2_{n-k}\bigg/(n-k)}}=\frac{\widehat y_i-y_i\bigg/\sigma\sqrt{\gamma_{ii}}}{\sqrt{(n-k)s^2\bigg/\sigma^2(n-k)}}=\frac{\widehat y_i-y_i}{s\sqrt{\gamma_{ii}}}\sim t(n-k) \tag{3.43} \end{equation}\]