Statistics I · Lecture 6
These slides are a free translation and adaptation from the slide deck for Estatística I by Prof. Sandra Custódio and Prof. Teresa Ferreira from the Lisbon Accounting and Business School, Polytechnic University of Lisbon.
A random variable is a function that will allow us to quantify (transform into a number) each outcome.
Random Variable
A random variable (r.v.) \(X\) is a function \(f:\Omega\rightarrow \Omega_X\subset \mathbb{R}\). \(\Omega_X\) is known as the support of the r.v. \(X\).
\[\omega\in\Omega \overset{X}{\rightarrow} X(\omega)\in\Omega_X\subset\mathbb{R}\]
\(X(\omega)\) is the image under \(X\) of the outcome \(\omega\)
Summarizing, a r.v. is a function that associates a real number to each outcome from \(\Omega\).
\(X\) is a discrete r.v. when:
In this case, \(\Omega_X=\{x_1, x_2, ... , x_n\}\) with \(n\in\mathbb{N}\) if \(\Omega_X\) is finite, and \(\Omega_X=\{x_1,x_2,...,x_n,...\}\) if it is countable infinite.
Let \(X\) be a discrete r.v. The pdf of \(X\) is a function \(f_X:\mathbb{R}\rightarrow\mathbb{R}\) such that:
\[f_X(x)=\left\{\begin{array}{cc}P(X=x) & ,\text{ if } x\in\Omega_X\\ 0 & ,\text{ if } x\in\mathbb{R}\setminus\Omega_X\end{array}\right.\]
Naturally, by construction the pdf satisfies the following properties:
| \(f_X(x)\geq 0 \quad \forall x\in\mathbb{R}\) |
| \(\sum_{x_i\in\Omega_X}P(X=x_i)=1\) |
The pdf gives the probability in a single point. The total probability is distributed among single points, \(x_i\). A reasonable representation of a pdf of a discrete r.v. could be:
| \(x\) | \(x_1\) | \(x_2\) | … | \(x_n\) | … |
|---|---|---|---|---|---|
| \(f(x)\) | \(p_1\) | \(p_2\) | … | \(p_n\) | … |
Where \(p_i=P(X=x_i)\)
Consider the discrete r.v. \(X\) with the following pdf:
| \(x\) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| \(f(x)\) | \(0.05\) | \(a\) | \(0.35\) | \(0.25\) | \(0.05\) |
We could define:
Our table is now:
| \(x\) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| \(f(x)\) | \(0.05\) | \(0.3\) | \(0.35\) | \(0.25\) | \(0.05\) |
What is \(P(X=2|X\leq 3)\)?
\[P(X=2|X\leq 3)=\frac{P(X=2 \cap X\leq 3)}{P(X\leq 3)}= \frac{P(X=2)}{P(X\leq 3)}\]
\[= \frac{f(2)}{f(0)+...+f(3)}=\frac{0.35}{0.95}=0.368\]
\(X\) is a continuous r.v. if:
Let \(X\) a continuous r.v.
There is a function \(f_X:\mathbb{R}\rightarrow\mathbb{R}\), the pdf of \(X\) such that:
Technically, from Measure Theory, we need an absolutely continuous r.v. to ensure the existence of a pdf. These issues are beyond the scope of this course. Just know that when we say continuous r.v. we mean absolutely continuous r.v.
Note that this pdf allows to compute the probability of events \(x\in(a,b]\):
\[P(a<X\leq b)=\int_a^b f_X(x)dx\]
Observe that if you would do \(X=a\) you would get the integral from \(a\) to \(a\), which makes \(dx=0\) and therefore the integral (and the probability) becomes 0.
Let \(X\) be a continuous r.v. with the following pdf:
\[ f(x)=\left\{\begin{array}{cc} \theta x^2 & , 0\leq x< 1\\ 0 & ,\mathbb{R}\setminus [0,1) \end{array}\right. \]
Support for \(X\): \(\Omega_X=[0,1)\)
\[\int_{-\infty}^{\infty}f(x)dx=1\Leftrightarrow\int_0^1\theta x^2dx=\left[\theta\frac{x^3}{3}\right]_{0}^1\]
\[\theta\frac{1}{3}-\theta\frac{0}{3}=1\Leftrightarrow \theta=3\]
Let \(X\) be a r.v. The distribution function \(F_X:\mathbb{R}\rightarrow[0,1]\), defined as:
\[F_X(x)=P(X\leq x)\]
\(F_X\) is unique.
With a discrete r.v.
| \(x\) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| \(f(x)\) | \(0.05\) | \(0.3\) | \(0.35\) | \(0.25\) | \(0.05\) |
\[F(x)=P(X\leq x)=\left\{ \begin{array}{cc} 0 & x<0 \\ 0.05 & 0\leq x < 1 \\ 0.05 + 0.3 = 0.35 & 1 \leq x < 2 \\ 0.35 + 0.35 = .7 & 2 \leq x < 3 \\ 0.7 + 0.25 = .95 & 3 \leq x < 4 \\ 1 & x\geq 4 \end{array} \right.\]
Let’s revisit our previous example:
\[P(X=2|X\leq 3)= \frac{P(X=2)}{P(X\leq 3)}=\] \[\frac{F(2)-F(2^-)}{F(3)}= \frac{0.7-0.35}{0.95}=0.368 \]
With a continuous r.v.:
\[F_X(x)=P(X\leq x)=\int_{-\infty}^{x} f_X(x)dx\]
The distribution function, \(F_X\) allows to compute the probability of \(\{X\in(a,b]\}\)
\[P(a<X\leq b)=\int_a^b f_X(x)dx=F_X(b)-F_X(a)\]
Consider the continuous r.v. defined previously, with the pdf:
\[ f(x)=\left\{ \begin{array}{cc} 3x^2 & , 0\leq x< 1\\ 0 & , \mathbb{R}\setminus [0,1) \end{array} \right. \]
Support for \(X\): \(\Omega_X=[0,1)\)
Distribution function (cdf):
\[ F(x)=P(X\leq x) = \int_{-\infty}^x f(t)dt = \left\{ \begin{array}{cc} 0 & , x<0\\ x^3 & ,0\leq x<1 \\ 1& ,x\geq 1 \end{array} \right. \]
Nonetheless the r.v. is discrete or continuous, \(F_X\) has the following properties:
\(F_X\) for \(X\) r.v. discrete
\(F_X\) for \(X\) r.v. continuous
For a discrete r.v.
\[P(X=x)=F_X(x)-F_X(x^-)\] \[+\downarrow \uparrow -\] \[F_X(x)=\sum_{x_i\leq x}P(X=x_i)\]
For a continuous r.v.
\[ f_X(x)=\left\{ \begin{array}{cc} F_X'(x) & ,x\in\mathbb{R} \text{ if }F_X'\text{ exists} \\ 0 & \text{, otherwise} \end{array} \right. \]
\[ Derivative \downarrow \uparrow Primitive\]
\[ F_X(x)=\int_{-\infty}^x f_X(t)dt\]
The pdf of \(X\), a continuous r.v. is not unique.
A continuous r.v. on \([0,6]\). Move \(x\): the shaded area under the pdf is exactly the height of the cdf at that point.
The cdf height is the pdf area. That is the whole content of \(F_X(x)=\int_{-\infty}^x f_X(t)dt\).
Watch the steepness too: where the pdf is tall the cdf climbs fastest, and once the pdf returns to zero the cdf flattens at 1 and stays there. That is the other direction, \(f_X(x)=F_X'(x)\).
The pdf height is not a probability: here it never exceeds \(1/3\), and for other variables it can exceed 1 entirely. Only the area is a probability.
If \(X\) is a continuous random variable, then \(P(X=a)\) equals:
A. \(f_X(a)\)
B. \(F_X(a)\)
C. \(0\)
D. \(1\)
✅ C. For a continuous r.v. the integral from \(a\) to \(a\) is zero. Probability lives on intervals, not on points.
The cumulative distribution function \(F_X\) is:
A. decreasing
B. always continuous
C. non decreasing, going from 0 to 1
D. the same as the pdf
✅ C. \(F_X(x)=P(X\leq x)\) can never decrease as \(x\) grows, and it goes from 0 to 1. For a discrete r.v. it is a step function, so not always continuous.
Let \(X\) be continuous with \(f(x)=2x\) for \(0\leq x<1\) and \(0\) otherwise.
Check that this is a valid pdf, find \(F_X(x)\), then compute \(P(X\leq 0.5)\) and \(P(0.2<X\leq 0.8)\).
Valid: \(f(x)=2x\geq 0\) on \([0,1)\) and \(\int_0^1 2x\,dx=\left[x^2\right]_0^1=1\) ✅
\(F_X(x)=\int_0^x 2t\,dt = x^2\) for \(0\leq x<1\) (and 0 below, 1 above).
\(P(X\leq 0.5)=0.5^2=0.25\)
\(P(0.2<X\leq 0.8)=F(0.8)-F(0.2)=0.64-0.04=0.60\)
We could describe the range of \(X\), a r.v., as a population, in the statistical sense, because it describes all the possible values it can take.
We can use numerical values to do so, which can represent dispersion or centrality of the data.
The expected value or mean, is a location parameter for our r.v.
Definition
The expected value or mean of a random variable \(X\) is:
Not all random variables have an expected value, it might be infinite.
Let \(X,Y\) rvs, and \(a,b\in\mathbb{R}\) scalars. Some properties of the mean
Let \(X\) be a discrete r.v. as in the previous example:
| \(x\) | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| \(f(x)\) | \(0.05\) | \(0.3\) | \(0.35\) | \(0.25\) | \(0.05\) |
Let \(g(X)=2(X-1)^2+3(X-1)-5\), find \(E[g(X)]\).
\[g(X)=2(X-1)^2+3(X-1)-5\] \[=2(X^2-2X+1)+3X-3-5\] \[=2X^2-4X+2+3X-8\] \[=2X^2-X-6\]
\[E[Y]=E[2X^2-X-6]\]
\[=2E[X^2]-E[X]-6\]
We only need to find \(E[X]\) and \(E[X^2]\) to obtain \(E[g(X)]\).
\[E[X]=\sum xP(X=x)\] \[ = 0\times .05 + 1 \times .3 + 2 \times .35 + 3 \times .25 + 4\times .05 = 1.95\]
\[E[X^2]=\sum x^2 P(X=x)\]
\[ = 0\times .05 + 1 \times .3 + 4 \times .35 + 9 \times .25 + 16\times .05 = 4.75\]
\[E[g(X)]=2\times 4.75 - 1.95 - 6 = 1.55\]
Recall our example for continuous r.v.s. \(X\): \[ f_X(x)=\left\{ \begin{array}{cc} 3x^2 & ,0\leq x < 1 \\ 0 & , \mathbb{R}\setminus[0,1) \end{array} \right. \]
Find \(E[g(X)]\) when \(g(X)=2(X-1)^2+3(X-1)-5\) We know already \(g(X)=2X^2-X-6\). Let’s focus on \(E[x]\) and \(E[X^2]\).
\[E[X]=\int_{-\infty}^{\infty} xf_X(x)dx = \int_0^1 x\times 3x^2 dx\] \[= \int_0^1 3x^3dx=\left[3\frac{x^4}{4}\right]_{0}^1=\frac{3}{4}=0.75\]
\[E[X^2]=\int_{-\infty}^{\infty} x^2f_X(x)dx = \int_0^1 x^2\times 3x^2 dx\] \[= \int_0^1 3x^4dx=\left[3\frac{x^5}{5}\right]_{0}^1=\frac{3}{5}=0.6\]
Finally,
\(E[g(x)]=2\times 0.6 - 0.75 - 6 = -5.55\)
The p-quantile, \(x_p\), of a r.v. \(X\) is a location parameter, with fixed value.
\(x_p\) is the value for \(x\in\Omega_X\) such that:
O what is the same, \(x\in\Omega_X\) such that \(F_X(x^-)\leq p \leq F_X(x)\)
\(x_p\) is an \(x\in\Omega_X\) such that \(F_X(x)=p\)
Let’s apply this for the examples we just used for the expected value.
Find the median (0.5-quantile) for \(X\)
\[ F(x)=P(X\leq x)=\left\{ \begin{array}{cc} 0 &, x\leq 0\\ 0.05 &, 0\leq x <1 \\ 0.35 &, 1\leq x < 2 \\ 0.7 &, 2 \leq x < 3 \\ 0.95 &, 3 \leq x < 4 \\ 1 &, x\geq 4 \end{array} \right. \]
For example \(F(2^-)=0.35\leq 0.5 \leq 0.7=F(2)\) and therefore \(X_{0.5}=Me = 2\). Given that \(E[X]=1.95<Me(X)=2\) the distribution is slightly negatively (or left) skewed.
\[F_X(x)=P(X\leq x)=\left\{ \begin{array}{cc} 0 & x < 0 \\ x^3 & 0\leq x <1 \\ 1 & x\geq 1 \end{array} \right. \]
Let’s find \(x\) such that \(F(x)=0.5\)
\(F(x)=0.5\Leftrightarrow x^3=0.5\Leftrightarrow x=\sqrt[3]{0.5}\approx0.7937\)
And therefore, \(x_{0.5}=Me=.7937\)
Let \(X\) be a r.v. The variance of \(X\), if it exists, is defined as:
\[V[X]=E\left[\left(X-E[X]\right)^2\right]\]
It can be shown, very easily, with some algebraic manipulation that \(V[X]=E\left[X^2\right]-\left(E[X]\right)^2\)
Remember that \(E[X]\equiv\mu_X\)
Usually we write \(V[X]\) as \(\sigma^2_X\).
Some properties for the variance:
If \(\sigma^2_X\) is the variance of \(X\), then the standard deviation is known as: \[\sigma_X=\sqrt{V[X]}\]
One characteristic of the standard deviation is that its units are the same as those of the random variable.
While the variance and standard deviation allow us to measure the dispersion of the data, we might want to have it relative to the mean (a \(\sigma_X=1\) can be a lot for \(X\) taking relatively low values, but negligible if we are talking in millions!)
For that we use the coefficient of variation:
\[C.V._X =\frac{\sigma_X}{\mu_X}\times 100\]
Some properties of the \(CV_X\)
Let’s compute \(\sigma^2\), \(\sigma\), and \(CV\) for our previous examples:
Values below \(50\%\) for \(CV\) allow us to see \(\mu\) as representative for the data. The lower, the closer the data to \(\mu\) and therefore the more representative it is.
For the continuous r.v. case:
For constants \(a\) and \(b\), \(V[aX+b]\) equals:
A. \(a^2V[X]\)
B. \(aV[X]+b\)
C. \(a^2V[X]+b^2\)
D. \(aV[X]\)
✅ A. Adding \(b\) shifts the whole distribution without changing its spread, so \(b\) drops out. Scaling by \(a\) scales the squared deviations by \(a^2\).
The coefficient of variation is useful because it:
A. is a pure number, measuring dispersion relative to the mean
B. keeps the units of \(X\)
C. is always below 1
D. equals the variance
✅ A. A \(\sigma\) of 1 is huge if the mean is 2, and negligible if the mean is a million. Dividing by the mean makes the comparison fair.
The number of faulty items \(X\) in a batch has \(f(0)=0.5\), \(f(1)=0.3\), \(f(2)=0.15\), \(f(3)=0.05\).
Compute \(E[X]\), \(V[X]\), \(\sigma_X\) and the coefficient of variation.
\(E[X]=0(0.5)+1(0.3)+2(0.15)+3(0.05)=0.75\)
\(E[X^2]=0(0.5)+1(0.3)+4(0.15)+9(0.05)=1.35\)
\(V[X]=1.35-0.75^2=0.7875\) · \(\sigma_X=\sqrt{0.7875}\approx 0.887\)
\(CV=\frac{0.887}{0.75}\times 100\approx 118\%\). Well above 50%, so the mean is not a representative summary of this distribution.
Problem set 2.2, Questions 1 to 9. They are all single random variables: pdf, cdf, quantiles and moments.
Next lecture we finish Topic 2 with random pairs, and then we review for the midterm.