counts_rigged = [10, 8, 12, 12, 16, 14]Univariate Statistics: Exercises
Set the seed to 10110.
Simulate a coin toss so that it displays both:
- the numeric outcome (
0or1) - the corresponding text (
"Heads"or"Tails")
Method 0:
- use an
ifstatement
Method 1:
- do not use an
ifstatement
Verify that both methods produce the same output.
Set the seed to 5005.
Generate five coin tosses. Use 3 different methods to print and store the result.
Method 0:
- store tosses as a single long string containing multiple occurences of
"Heads"and"Tails".
Method 1:
- store tosses as a list containing
"Heads"and"Tails"as its elements.
Method 2:
- store tosses as a list containing
0and1as its elements.
For each method, write code to determine whether the outcome of the best-of-five match is already decided after:
- 3 tosses
- 4 tosses
- 5 tosses
Print the winner and the number of tosses required.
Set the seed to 50052000. Using your code from exercise 5005, run 2000 simulations and compute the frequency that the outcome of the best-of-five match is already decided after:
- 3 tosses
- 4 tosses
- 5 tosses
Set the seed to 4242.
Create a list containing four d6 rolls.
Replace only the 2-th element of the list by a new d6 roll.
Print:
- the original list
- the modified list
- the difference between the two lists
Set the seed to 5555.
Roll a d6 repeatedly.
Keep only rolls equal to 5.
Stop after collecting four 5 results.
Print the total number of rolls required.
Set the seed to 5555200.
Create an empty list num_rolls_four_5.
Repeat exercise 5555 exactly 200 times.
After each experiment, append the number of rolls required to num_rolls_four_5
Print the last 12 elements of the list num_rolls_four_5, and its length.
Set the seed to 5555201.
compute manually the following descriptive statistics for num_rolls_four_5
- mean
- minimum
- maximum
- median
- first quartile
- third quartile
- variance
Print all results.
.describe()
Set the seed to 5555202.
Create a dataframe containing a single data series num_rolls_four_5. Use .describe() to check your calculations from Exercise 5555201.
Set the seed to 5555203.
Create a histogram of num_rolls_four_5.
Add vertical lines showing:
- the mean
- the median
- the 90-th percentile
Set the seed to 6666.
Roll a d6 repeatedly.
Stop only when four consecutive rolls equal:
6 6 6 6
Print the total number of rolls required.
Set the seed to 6666200.
Repeat Exercise 6666 exactly 200 times.
Store the waiting times in: num_rolls_four_6
Print:
- the minimum waiting time
- the average waiting time
- the maximum waiting time
Set the seed to 31415.
Roll a d6 repeatedly.
Stop only when:
3 1 4 1 5
appears consecutively.
Print the total number of rolls required.
Set the seed to 31415200.
Repeat Exercise 31415 exactly 200 times.
Store the waiting times in:
num_rolls_pi
Construct a dataframe containing:
- summary statistics for
num_rolls_pi - summary statistics for
num_rolls_four_6
Print the dataframe.
Which experiment has the larger average waiting time?
Set the seed to 29365.
Assume:
- a group contains exactly
29students - each birthday is equally likely
- there are
365possible birthdays 29 Februarydoes not exist
Simulate 4000 groups.
For each group determine whether any of the following occurs:
- no two students share a birthday
- at least two different student pairs share a birthday
- at least one birthday is shared by three or more students
Using the simulations, compute: - frequency of “no shared birthdays” - frequency of “multiple pairs share a birthday” - probability of “at least one triple-or-more match”
Set the seed to 2621121.
Create 4000 realizations of:
2d6+2, i.e. the result of rolling 2 d6, adding the faces, and adding 2
and 4000 realizations of:
1d12+1, i.e. the result of rolling 1 d12, and adding 1
Store the results in:
rolls_2d6_plus_2
rolls_1d12_plus_1
Put them both in the same pandas dataframe.
Compute and print:
- minimum
- maximum
- mean
- variance
for both variables.
A “successful roll” is defined as a roll where the realization is 10 or above. Between 2d6+2 and 1d12+1, which has the higher frequency of successful rolls?
Set the seed to 10141212161400.
Using:
define:
H0 : the d6 is rigged
H1 : the d6 is fair
Generate the sampling distribution of the sample mean under H0.
Display the histogram.
Set the seed to 10141212161401.
Using the reversed hypotheses,
construct a test with:
alpha = 0.10
sample_size = 50Print the rejection cutoff.
Set the seed to 10141212161402.
Using the reversed hypotheses,
simulate 20 unknown d6 samples.
For each sample print:
- sample mean
- p-value
- reject / fail to reject H0
Set the seed to 10141212161403.
Estimate:
- type I error probability
- type II error probability
- power
for the reversed test.
Print the results in a dataframe.
Set the seed to 101412121614.
Use:
counts_rigged = [10, 8, 12, 12, 16, 14]
sample_size = 50
n_simulations=400Generate:
n_simulationssample means from the fair d6n_simulationssample means from the rigged d6
Display both distributions on the same graph.
Set the seed to 10141212161499.
Use a different variant of a rigged d6, by changing interval_widths_rigged to [14, 10, 12, 12, 8, 16]. Consider sample_size = 50, alpha = 0.10, n_simulations = 3000
The exercise contrasts three possible test statistics to test hypothesis \(H_0\): the d6 is fair. - sample mean - sample variance - log likelihood ratio
For each test statistic, - generate n_simulations realizations of the test statistic, - plot the distribution of the test statistic under \(H_0\), - apply a jitter, - choose the cutoff - simulate the test n_simulations times under \(H_1\): the d6 is rigged
Construct a dataframe with one row per test statistic, and with colums for: - cutoff - empirical alpha - count of simulations where you reject \(H_0\) - count of simulations where you fail to reject \(H_0\) - empirical power of the test statistic
Sort the dataframe by decreasing empirical power and print the dataframe.
Source: Cyganowski, S., Kloeden, P. E., and Ombach, J. (2002), From Elementary Probability to Stochastic Differential Equations with Maple, Springer.
A population of 1200 individuals is screened for a rare virus. Each individual is infected independently with probability p = 0.01. Testing is expensive. To reduce costs, the testing center practices pooled testing:
- divide the population into groups of 30 individuals
- mix all blood samples within a group
- perform a single pooled test on the mixture
- if the pooled test is negative, stop
- if the pooled test is positive, test every individual in that group separately
Set the seed to 1200110030.
Simulate the testing campaign once.
Print the total number of tests performed.
Based on Exercise 1200110030.
set the seed to 1200110030400.
Simulate the testing campaign n_simulations = 400 times and store the number of tests performed.
Print: - the minimum number of tests - the average number of tests - the maximum number of tests
Set the seed to 12001100145400. Keep using p = 0.01 and n_simulations = 400 Test different group test sizes between 1 and 45.
Construct a dataframe containing: - group size - the minimum number of tests - the average number of tests - the maximum number of tests
Print the group size with the lowest average number of tests.
Set the seed to 120051000145400.
Repeat Exercise 12001100145400, except using: p = 0.005 instead of p = 0.01.
Print: - the optimal group size - the corresponding average number of tests