Statistics I · Lecture 4
These slides are a free translation and adaptation from the slide deck for Estatística I by Prof. Sandra Custódio and Prof. Teresa Ferreira from the Lisbon Accounting and Business School, Polytechnic University of Lisbon.
Consider the following scenario:
💼 Investor: What’s the probability this startup will succeed?
📊 Analyst: Hard to say, every startup is different.
💼 Investor: But if you had to guess, based on similar cases?
📊 Analyst: Maybe 1 in 3 succeed under these conditions.
💼 Investor: So, would you bet on it?
📊 Analyst: Yes, I would.
💼 Investor: Even if the odds aren’t great?
📊 Analyst: I believe this one has what it takes.
Here we can define probability in terms of frequency of occurrence, i.e. as a percentage of successes in a moderately large number of similar situations.
This is the most natural and traditional way of thinking about probability.
But, what if this company belongs to a completely novel market sector?
There might be situations where the frequency concept is not adequate, because it might refer to a one-time event. These are subjective beliefs.
A company is recruiting a new CEO, and a board member says:
“I believe there’s a 90% chance that our chosen candidate will be an effective CEO.”
It might seem easy to disregard the second case as unscientific or useless. However, many times people need to make decisions under uncertainty with not enough data (or no data at all!) about previous realizations of the specific event.
Beliefs allow the decision maker to, well, make some decision, at least consistently.
Uncertainty
A set is a collection of objects, which are elements of the set.
Definition
Let \(S\) represent a set, and \(s\) an element of that set, we write \(s\in S\) to mean \(s\) belongs to \(S\).
If \(s\) does not belong to \(S\), we write \(s\notin S\).
Definition
If a set \(S\) does not have any element, then it is the empty set, denoted by \(\emptyset\).
There are several ways to specify a set.
By extension, or as a list:
If a set \(S\) has a finite number of elements (\(x_i\in S\)) we can write it like this: \[S=\{x_1, x_2, ..., x_n\}\]
If a set \(S\) has an infinite (but countable) elements (\(x_i\in S\)) we can write it like: \[S=\{x_1, x_2, ...\}\]
By describing the property (\(P\)) that \(x\) must satisfy to be included in \(S\): \[S= \{x|x \text{ satisfies } P\}\] in this case \(|\) reads as such that. For example \[S=\{x\in\mathbb{R}\mid x\geq 0\}\] to describe the non-negative real numbers.
This example is special, as the positive real numbers cannot be written down as a a list. In this case the interval \([0,\infty)\) is an uncountable set.
Definition
If \(\forall x\in S\) it is also true that \(x\in T\), then we say that \(S\) is a subset of \(T\), and we write it like \(S\subseteq T\).
Definition
If \(S\subseteq T\) and at the same time \(T\subseteq S\) then we say that \(S\) and \(T\) are equal, and we write it \(S=T\).
Definition
The universal set \(\Omega\) is the set that contains all objects that could conceivably be of interest in a particular context.
By definition, any set \(S\) must be a subset of \(\Omega\).
The universal set is important because it defines the scope of our analysis. Say we are studying the performance of students of Statistics I in 2026.
The cars parked outside our institution do not belong to the universal set, because they are not relevant for our purpose. Only students of Statistics I in 2026 belong to the universal set.
Definition
The complement of a set \(S\), with respect to \(\Omega\), is the set \(\{x\in\Omega| x\notin S\}\), that is, all the relevant elements that do not belong in \(S\). We denote it as \(S^c\).
Corollary: It is easy to see that \(\Omega^c=\emptyset\).
Definition
The union of two sets \(S\) and \(T\) is the set of all elements that belong to \(S\) or \(T\) (or both), and is denoted by \(S\cup T\). \[S\cup T=\{x\in\Omega | x\in S\ \vee\ x\in T\}\]
Definition
The intersection of two sets \(S\) and \(T\) is the set of all elements that belong to \(S\) and \(T\), and is denoted by \(S\cap T\). \[S\cap T=\{x\in\Omega | x\in S\wedge x\in T\}\]
Note that \(\vee\) stands for or, and \(\wedge\) stands for and.
Sometimes we might need to consider the union or intersection of many sets, and for that we can use a notation similar to the one we used for summations:
\[\bigcup_{n=1}^\infty S_n = S_1 \cup S_2 \cup ... = \{x\in\Omega | x \in S_n \text{ for some } n\}\]
\[\bigcap_{n=1}^\infty S_n = S_1 \cap S_2 \cap ... = \{x\in\Omega | x \in S_n \text{ for every } n\}\]
Definition
Two sets (say \(S\) and \(T\)) are said to be disjoint if \(S\cap T=\emptyset\).
More generally, a collection of sets \(S_n\) is disjoint if \(S_i\) and \(S_j\) are disjoint when \(i\neq j\).
Definition
A collection of sets is said to be a partition of a set \(S\) if the sets in the collection are:
Disjoint
Their union is \(S\)
The set of all subsets of \(\Omega\) is called the power set, denoted \(\mathcal{P}(\Omega)\). A partition of \(\Omega\) is then a collection \(\{A_i\}\subseteq\mathcal{P}(\Omega)\).
The number of elements of a set \(S\) is known as its cardinality and it is denoted as \(\# S\). \(\# S\) satisfies:
If we have two sets, \(S\) and \(T\), we define the operation set minus as \(\setminus\) the set that contains all elements of \(S\) that do not belong to \(T\)
\[S\setminus T = \{x\in S| x\notin T\}\]
Two sets \(S\) and \(T\) are disjoint when:
A. \(S\cup T=\emptyset\)
B. \(S=T\)
C. \(S\cap T=\emptyset\)
D. \(S\subseteq T\)
✅ C. Disjoint means they share no element at all, so their intersection is empty. Their union is empty only if both sets are.
The set \(S\setminus T\) contains exactly the elements that:
A. belong to both \(S\) and \(T\)
B. belong to \(T\) but not to \(S\)
C. belong to neither
D. belong to \(S\) but not to \(T\)
✅ D. \(S\setminus T=\{x\in S\,|\,x\notin T\}\). Note the operation is not symmetric: \(S\setminus T\) and \(T\setminus S\) are different sets.
Let \(\Omega=\{1,2,3,4,5,6\}\), \(A=\{1\}\), \(B=\{3,6\}\) and \(C=\{2,4,6\}\).
Write down \(A\cup B\), \(B\cap C\), \((B\cup C)^c\) and \(C\setminus B\).
\(A\cup B=\{1,3,6\}\)
\(B\cap C=\{6\}\)
\(B\cup C=\{2,3,4,6\}\), so \((B\cup C)^c=\{1,5\}\)
\(C\setminus B=\{2,4\}\)
In probability, \(\Omega\), the universal set, is a non-empty set that contains all possible outcomes of an experiment. Each outcome is represented by \(\omega\), and obviously \(\omega\in\Omega\).
The sample space (\(\Omega\)) can be:
Consider the experiment of throwing a die 🎲 and noting the number shown on side facing upwards.
Consider now the random experiment of measuring the life expectancy of a lamp 💡, measured in hours.
Remember, \(\Omega\) must include all possible outcomes from your experiment! Even then ones that seem ludicrous.
Definition
A subset \(A\) of the sample space \(\Omega\) is called an event. \[A\subseteq \Omega\]
By definition then \(\Omega\) is also an event.
Definition
We call the realization of an event \(A\) if, after an experiment, outcome \(\omega\) is realized, and \(\omega \in A\).
Let’s go back to our experiment with the 🎲
The sample space is: \(\Omega=\{1,2,3,4,5,6\}\)
Within this space, we can define the following events:
Now let’s revisit the example of our 💡
The sample space is: \(\Omega=\{x\in\mathbb{R}|x\geq 0\}\)
In this space, we can define the following events:
Definitions
An elementary event is any event that contains a single element (i.e. \(\# A = 1\))
An impossible event is an event with no outcome, (i.e. \(\# A=0\)). As a consequence, an impossible event coincides with the empty set \(\emptyset\).
A certain event is indeed the event \(\Omega\), as for any outcome we obtain \(\omega\), this outcome belongs to the sample space \(\Omega\) by definition.
Consider two events \(A\) and \(B\) both subsets of \(\Omega\)
\(A^c\) contains all the outcomes that are not in \(A\). \(A^c\) is the event of not \(A\).
If \(A\subseteq B\), then an outcome that realizes event \(A\) (\(\omega\in A\)), also realizes \(B\), as \(A\subseteq B\Rightarrow \omega\in B\) as well. \(A\Rightarrow B\)
For \(A\cup B\) to happen, we need \(\omega \in A\) or \(\omega \in B\), which means that \(A\) happens, or \(B\) happens, or both happen simultaneously.
For \(A\cap B\) to happen, we need \(\omega \in A\) and \(\omega \in B\), which means that \(A\) and \(B\) happen simultaneously.
\(A\) and \(B\) are incompatible if \(A\cap B=\emptyset\), i.e. if an outcome is in one set, it cannot be in another, for example it cannot be that \(A\) and \(A^c\) happen simultaneously!
Consider two events \(A\) and \(B\) both subsets of \(\Omega\)
Consider the sample space \(\Omega =\{1,2,3,4,5,6\}\), from our 🎲 case.
Define the events: \[A=\{1\},\ B=\{3,6\},\ C=\{2,4,6\},\ D=\{4,5,6\}\]
Let’s define the following events in \(\Omega\)
Besides the concepts we already saw of frequency and subjectivity for probability, there was an older, called “classic” one. This one was introduced by Pierre-Simon Laplace in 1812.
Laplace or Classic interpretation of probability
Let \(A\) be an event defined over a finite \(\Omega\). The probability of event \(A\) is defined as:
\[P(A)=\frac{\# A}{\# \Omega}\]
Consider now an experiment throwing two dice 🎲 🎲
The problem with this interpretation, is that we cannot use it, or it becomes meaningless, it when \(\Omega\) is uncountable or infinite. Also, what if the outcomes are not equally likely? (i.e. if the dice are not fair?)
This is today still the dominant interpretation of probability.
In this case, what we want is to observe several independent repetitions of the experiment. After a while, some statistical regularity begins to emerge.
Logically, if you run an experiment, and are interested in the probability of event \(A\), then the events you are registering are \(A\) and \(A^c\) or not \(A\).
Every time you run your experiment, you count when you get an \(A\) and when you observe an \(A^c\) event. Obviously, the total number of experiments is how many times you observed \(A\) and how many times you observed \(A^c\).
Experiment: Draw a random number in the interval \([0,1]\). \(A\) denotes \(x<0.4\).
| \(A\) | \(A^c\) | \(N\) | \(P(A)\) |
|---|---|---|---|
| 0 | 1 | 1 | 0 |
| 2 | 8 | 10 | 0.2 |
| 17 | 33 | 50 | 0.34 |
| 39 | 61 | 100 | 0.39 |
| 217 | 283 | 500 | 0.434 |
| 802 | 1198 | 2000 | 0.401 |
As you can see, the more experiments we run, the more stabilized the ratio of occurrences for \(A\) over the total number of experiments. More generally:
\[P(A)=\lim_{N\rightarrow \infty}\frac{A\text{ occurrences}}{N \text{- Number of Experiments}}\]
This is the relative frequency of \(A\) in \(N\) experiments: \(f_A\)
Not always possible to repeat that many times the experiment in the same conditions.
It seems that the probability that the random number between 0 and 1 is below 0.4 is approximately 40%. The more experiments we run, the closer our relative frequency is to that number.
\[P(A)\underset{N \rightarrow \infty}{\rightarrow} 0.4\]
By the way, we will see later that theoretically, indeed \(P(A)=0.4\)
You roll two fair dice. The probability that the sum equals 7 is:
A. \(\frac{1}{12}\)
B. \(\frac{1}{6}\)
C. \(\frac{7}{36}\)
D. \(\frac{1}{9}\)
✅ B. There are 6 favourable outcomes out of 36: \((1,6),(2,5),(3,4),(4,3),(5,2),(6,1)\), so \(P=\frac{6}{36}=\frac{1}{6}\).
Under the frequency interpretation, \(P(A)\) is:
A. a subjective degree of belief
B. always \(\frac{\# A}{\#\Omega}\)
C. the limit of the relative frequency of \(A\) as the number of experiments grows
D. always \(0.5\)
✅ C. The frequency interpretation repeats the experiment many times and takes \(P(A)=\lim_{N\to\infty}\frac{\text{occurrences of }A}{N}\). The \(\frac{\# A}{\#\Omega}\) formula is the classical definition instead, and it needs equally likely outcomes.
Roll two fair dice. Let \(A\) be “both dice show the same number” and \(B\) be “the sum is at least 10”.
Compute \(P(A)\), \(P(B)\) and \(P(A\cap B)\).
\(\#\Omega=36\).
\(A=\{(1,1),...,(6,6)\}\), so \(P(A)=\frac{6}{36}=\frac{1}{6}\).
\(B=\{(4,6),(5,5),(6,4),(5,6),(6,5),(6,6)\}\), so \(P(B)=\frac{6}{36}=\frac{1}{6}\).
\(A\cap B=\{(5,5),(6,6)\}\), so \(P(A\cap B)=\frac{2}{36}=\frac{1}{18}\).
If there is time left, two quick ones. Difficulty goes 🟢 warm-up, 🟡 standard, 🟠 exam level, 🔴 stretch.
Consider the following events:
A. “get a 7 when rolling a cubic die 🎲”
B. “Spain wins the next World Cup ⚽”
C. “rain in London ☔”
Which is correct?
✅ 3. Nobody can repeat the next World Cup many times, nor list equally likely outcomes. What is left is a belief.
When computing a probability from the analysis of all possible (equally likely) outcomes, we are using:
A. the frequentist definition of probability
B. the classic definition of probability
C. the subjective definition of probability
D. the axiomatic definition of probability
✅ B. This is Laplace: favourable cases over possible cases.
Problem set 2.1, Questions 1 and 2, to check the ideas of today.