Univariate Statistics: Exercises

TipPractice exercise 10110 — Pure coding — Same task, different implementation

Set the seed to 10110.

Simulate a coin toss so that it displays both:

  • the numeric outcome (0 or 1)
  • the corresponding text ("Heads" or "Tails")

Method 0:

  • use an if statement

Method 1:

  • do not use an if statement

Verify that both methods produce the same output.

TipPractice exercise 5005 — Pure coding — Choosing a useful representation

Set the seed to 5005.

Generate five coin tosses. Use 3 different methods to print and store the result.

Method 0:

  • store tosses as a single long string containing multiple occurences of "Heads" and "Tails".

Method 1:

  • store tosses as a list containing "Heads" and "Tails" as its elements.

Method 2:

  • store tosses as a list containing 0 and 1 as its elements.

For each method, write code to determine whether the outcome of the best-of-five match is already decided after:

  • 3 tosses
  • 4 tosses
  • 5 tosses

Print the winner and the number of tosses required.

TipPractice exercise 50052000 — Best-of-five

Set the seed to 50052000. Using your code from exercise 5005, run 2000 simulations and compute the frequency that the outcome of the best-of-five match is already decided after:

  • 3 tosses
  • 4 tosses
  • 5 tosses
TipPractice exercise 4242 — Coding comparison — Append versus indexing

Set the seed to 4242.

Create a list containing four d6 rolls.

Replace only the 2-th element of the list by a new d6 roll.

Print:

  • the original list
  • the modified list
  • the difference between the two lists
TipPractice exercise 5555 — Simulation — Collect four 5s

Set the seed to 5555.

Roll a d6 repeatedly.

Keep only rolls equal to 5.

Stop after collecting four 5 results.

Print the total number of rolls required.

TipPractice exercise 5555200 — Simulation + statistics — Repeat the experiment 200 times

Set the seed to 5555200.

Create an empty list num_rolls_four_5.

Repeat exercise 5555 exactly 200 times.

After each experiment, append the number of rolls required to num_rolls_four_5

Print the last 12 elements of the list num_rolls_four_5, and its length.

TipPractice exercise 5555201 — Simulation + statistics — Compute descriptive statistics manually

Set the seed to 5555201.

compute manually the following descriptive statistics for num_rolls_four_5

  • mean
  • minimum
  • maximum
  • median
  • first quartile
  • third quartile
  • variance

Print all results.

TipPractice exercise 5555202 — Simulation + statistics — Verify using .describe()

Set the seed to 5555202.

Create a dataframe containing a single data series num_rolls_four_5. Use .describe() to check your calculations from Exercise 5555201.

TipPractice exercise 5555203 — Visualization — Histogram of waiting times

Set the seed to 5555203.

Create a histogram of num_rolls_four_5.

Add vertical lines showing:

  • the mean
  • the median
  • the 90-th percentile
TipPractice exercise 6666 — Simulation — Four consecutive sixes

Set the seed to 6666.

Roll a d6 repeatedly.

Stop only when four consecutive rolls equal:

6 6 6 6

Print the total number of rolls required.

TipPractice exercise 6666200 — Simulation + statistics — Repeat four consecutive sixes

Set the seed to 6666200.

Repeat Exercise 6666 exactly 200 times.

Store the waiting times in: num_rolls_four_6

Print:

  • the minimum waiting time
  • the average waiting time
  • the maximum waiting time
TipPractice exercise 31415 — Simulation — Waiting for π

Set the seed to 31415.

Roll a d6 repeatedly.

Stop only when:

3 1 4 1 5

appears consecutively.

Print the total number of rolls required.

TipPractice exercise 31415200 — Simulation + statistics — Comparing two waiting-time problems

Set the seed to 31415200.

Repeat Exercise 31415 exactly 200 times.

Store the waiting times in:

num_rolls_pi

Construct a dataframe containing:

  • summary statistics for num_rolls_pi
  • summary statistics for num_rolls_four_6

Print the dataframe.

Which experiment has the larger average waiting time?

TipPractice exercise 29365 — Simulation — The birthday problem

Set the seed to 29365.

Assume:

  • a group contains exactly 29 students
  • each birthday is equally likely
  • there are 365 possible birthdays
  • 29 February does not exist

Simulate 4000 groups.

For each group determine whether any of the following occurs:

  • no two students share a birthday
  • at least two different student pairs share a birthday
  • at least one birthday is shared by three or more students

Using the simulations, compute: - frequency of “no shared birthdays” - frequency of “multiple pairs share a birthday” - probability of “at least one triple-or-more match”

TipPractice exercise 2621121 — Simulation + statistics — Comparing 2d6+2 and 1d12+1

Set the seed to 2621121.

Create 4000 realizations of:

2d6+2, i.e. the result of rolling 2 d6, adding the faces, and adding 2

and 4000 realizations of:

1d12+1, i.e. the result of rolling 1 d12, and adding 1

Store the results in:

rolls_2d6_plus_2

rolls_1d12_plus_1

Put them both in the same pandas dataframe.

Compute and print:

  • minimum
  • maximum
  • mean
  • variance

for both variables.

A “successful roll” is defined as a roll where the realization is 10 or above. Between 2d6+2 and 1d12+1, which has the higher frequency of successful rolls?

TipPractice exercise 10141212161400 — Hypothesis testing — Reverse the hypotheses

Set the seed to 10141212161400.

Using:

counts_rigged = [10, 8, 12, 12, 16, 14]

define:

H0 : the d6 is rigged
H1 : the d6 is fair

Generate the sampling distribution of the sample mean under H0.

Display the histogram.

TipPractice exercise 10141212161401 — Hypothesis testing — Construct a new rejection rule

Set the seed to 10141212161401.

Using the reversed hypotheses,

construct a test with:

alpha = 0.10
sample_size = 50

Print the rejection cutoff.

TipPractice exercise 10141212161402 — Hypothesis testing — Compute p-values under the new H0

Set the seed to 10141212161402.

Using the reversed hypotheses,

simulate 20 unknown d6 samples.

For each sample print:

  • sample mean
  • p-value
  • reject / fail to reject H0
TipPractice exercise 10141212161403 — Statistical power — Consequences of reversing H0 and H1

Set the seed to 10141212161403.

Estimate:

  • type I error probability
  • type II error probability
  • power

for the reversed test.

Print the results in a dataframe.

TipPractice exercise 101412121614 — Sampling distributions — When the mean starts to fail

Set the seed to 101412121614.

Use:

counts_rigged = [10, 8, 12, 12, 16, 14]
sample_size = 50
n_simulations=400

Generate:

  • n_simulations sample means from the fair d6
  • n_simulations sample means from the rigged d6

Display both distributions on the same graph.

TipCapstone exercise 10141212161499 — Capstone — Mean, variance and log likelihood ratio

Set the seed to 10141212161499.

Use a different variant of a rigged d6, by changing interval_widths_rigged to [14, 10, 12, 12, 8, 16]. Consider sample_size = 50, alpha = 0.10, n_simulations = 3000

The exercise contrasts three possible test statistics to test hypothesis \(H_0\): the d6 is fair. - sample mean - sample variance - log likelihood ratio

For each test statistic, - generate n_simulations realizations of the test statistic, - plot the distribution of the test statistic under \(H_0\), - apply a jitter, - choose the cutoff - simulate the test n_simulations times under \(H_1\): the d6 is rigged

Construct a dataframe with one row per test statistic, and with colums for: - cutoff - empirical alpha - count of simulations where you reject \(H_0\) - count of simulations where you fail to reject \(H_0\) - empirical power of the test statistic

Sort the dataframe by decreasing empirical power and print the dataframe.

TipPractice exercise 1200110030 — Simulation and optimization — Pooled virus testing

Source: Cyganowski, S., Kloeden, P. E., and Ombach, J. (2002), From Elementary Probability to Stochastic Differential Equations with Maple, Springer.

A population of 1200 individuals is screened for a rare virus. Each individual is infected independently with probability p = 0.01. Testing is expensive. To reduce costs, the testing center practices pooled testing:

  1. divide the population into groups of 30 individuals
  2. mix all blood samples within a group
  3. perform a single pooled test on the mixture
  4. if the pooled test is negative, stop
  5. if the pooled test is positive, test every individual in that group separately

Set the seed to 1200110030.

Simulate the testing campaign once.

Print the total number of tests performed.

TipPractice exercise 1200110030400 — Simulation and optimization — Repeated pooled testing

Based on Exercise 1200110030.

set the seed to 1200110030400.

Simulate the testing campaign n_simulations = 400 times and store the number of tests performed.

Print: - the minimum number of tests - the average number of tests - the maximum number of tests

TipCapstone exercise 12001100145400 — Simulation and optimization — Searching for the optimal group size

Set the seed to 12001100145400. Keep using p = 0.01 and n_simulations = 400 Test different group test sizes between 1 and 45.

Construct a dataframe containing: - group size - the minimum number of tests - the average number of tests - the maximum number of tests

Print the group size with the lowest average number of tests.

TipPractice exercise 120051000145400 — Simulation and optimization — A rarer virus

Set the seed to 120051000145400.

Repeat Exercise 12001100145400, except using: p = 0.005 instead of p = 0.01.

Print: - the optimal group size - the corresponding average number of tests