Danho
ZIMSEC A Level · 6046/2

Statistics Paper 2 Specimen 2026

Questions
63
Total marks
152
Time allowed
180 min
Syllabus code
6046/2

Sit this paper online

Questions
63
Pass mark
38
Sit this paper

Answer every question in the printed order, get marked at the end, then see the answers.

The questions

Question 101

[3 marks]continuous random variables
A continuous random variable XX has probability density function f(x)=0.1x+kf(x) = 0.1x + k for 4≤x≤64 \leq x \leq 6, f(x)=0.3f(x) = 0.3 for 6≤x≤86 \leq x \leq 8, and f(x)=0f(x) = 0 otherwise. Find the value of the constant kk.

Answer this when you sit the paper.

Question 102

[3 marks]continuous random variables
A continuous random variable XX has probability density function f(x)=0.1x−0.3f(x) = 0.1x - 0.3 for 4≤x≤64 \leq x \leq 6, f(x)=0.3f(x) = 0.3 for 6≤x≤86 \leq x \leq 8, and f(x)=0f(x) = 0 otherwise. Find P(5≤X≤7)P(5 \leq X \leq 7).

Answer this when you sit the paper.

Question 201

[1 marks]data presentation

Examination marks are displayed as below, with 1 alongside a leaf 3 standing for the mark 13 %.

StemLeaves
13
26
31
41 3
50 2 6 8
61 2 2 2 7
70 3 4 5 5 8 9
80 3 4 4 8
92 7 7 8

key: 4|1 means 41 %

What is the name of this type of display?

  1. AA stem and leaf diagram, which splits each value into a stem and a leaf
  2. BA histogram, whose bar areas are proportional to the class frequencies
  3. CA cumulative frequency curve, which plots running totals against class boundaries
  4. DA box and whisker plot, which shows the median, the quartiles and the extremes

Question 202

[3 marks]data presentation

The marks obtained by candidates in a mathematics examination are shown below, where a stem of 1 with a leaf of 3 means 13 %.

StemLeaves
13
26
31
41 3
50 2 6 8
61 2 2 2 7
70 3 4 5 5 8 9
80 3 4 4 8
92 7 7 8

key: 4|1 means 41 %

Calculate the range of the marks.

Answer this when you sit the paper.

Question 203

[2 marks]data presentation
For a set of 30 examination marks the mean is 67.0 %, the median is 71.5 %, the lower quartile is 55 % and the upper quartile is 83.25 %. What does this say about the skewness of the distribution?
  1. AIt is symmetrical, because the mean and the median both lie inside the interquartile range, which is what symmetry requires of a distribution
  2. BIt is positively skewed, because the upper quartile is further from zero than the lower quartile is
  3. CIt is negatively skewed, since the mean lies below the median and the lower half of the middle 50 % is the wider one
  4. DIt is positively skewed, since the mean lies below the median and the upper tail must therefore be the longer one

Question 204

[2 marks]data presentation
Examination marks are recorded in a stem and leaf diagram rather than in a grouped frequency table. Which of the following is an advantage of the stem and leaf diagram?
  1. AIt shows the shape of the distribution while hiding the individual values, which keeps the display uncluttered
  2. BIt keeps every original data value, so the median, the range and the mode can all be read straight off it
  3. CIt removes the need to order the data before the median can be found from the display
  4. DIt replaces the raw values by class mid-points, which makes the mean quicker to calculate

Question 301

[2 marks]discrete random variables
A fair spinner lands on one of the six values 1, 2, 3, 4, 5 and 6, each equally likely. Find the probability that the spinner lands on a prime value.

Answer this when you sit the paper.

Question 302

[3 marks]discrete random variables

A fair spinner lands on one of six values, each equally likely, and each value pays the prize shown.

value123456
prize in $2264106

Find the probability that the spinner lands on a value that gives a prize of not less than $4.

Answer this when you sit the paper.

Question 303

[3 marks]discrete random variables

A fair spinner lands on one of six values, each equally likely, and each value pays the prize shown.

value123456
prize in $2264106

Calculate the expected prize for a single game.

Answer this when you sit the paper.

Question 401

[2 marks]normal and binomial distributions
The masses of letters are normally distributed with mean 15 g, and 92 % of them lie within 10 g of the mean. Since 92 % lies between the two symmetric limits, 4 % lies in each tail. State the value of zz for which P(Z<z)=0.96P(Z < z) = 0.96, to 3 decimal places.

Answer this when you sit the paper.

Question 402

[2 marks]normal and binomial distributions
The masses of letters posted by a school are normally distributed with mean 15 g, and the masses of 92 % of the letters are within 10 g of the mean. Given that P(Z<1.751)=0.96P(Z < 1.751) = 0.96, find the standard deviation of the masses, in grams to 2 decimal places.

Answer this when you sit the paper.

Question 403

[2 marks]normal and binomial distributions
Each letter posted by a school has probability 0.92 of having a mass within 10 g of the mean, independently of the others. A random sample of 8 letters is taken. Which distribution does the number of letters in the sample whose mass is within 10 g of the mean follow?
  1. AA Poisson distribution with mean 8×0.92=7.368 \times 0.92 = 7.36, since the letters are posted at random and independently
  2. BA binomial distribution with 8 trials and probability of success 0.92
  3. CA normal distribution with mean 8 and standard deviation 0.92
  4. DA binomial distribution with 8 trials and probability of success 0.08

Question 404

[3 marks]normal and binomial distributions
The number of letters within 10 g of the mean, out of a random sample of 8, follows the binomial distribution B(8,0.92)B(8, 0.92). Find the probability that at least 2 of the 8 letters are within 10 g of the mean, correct to 3 decimal places.
  1. A0.083, which is 8(0.92)(0.08)78(0.92)(0.08)^7 doubled to allow for either of the two letters
  2. B0.632, since at least 2 letters out of a sample of 8 is only a little under two thirds likely
  3. C1.000, since the chance of 0 or 1 such letter is about 1.6×10−71.6 \times 10^{-7}
  4. D0.001, since two or more successes out of only 8 trials is very unlikely

Question 501

[2 marks]sampling
In statistics, what is meant by a random sample?
  1. AA sample chosen by taking every kkth member of the population from a list, after a start is fixed
  2. BA sample chosen so that it contains the same proportion of each group as the population does
  3. CA sample chosen without any plan at all, so that the sampler simply takes whichever members of the population happen to come to hand first, in whatever order they appear
  4. DA sample chosen so that every member of the population has an equal chance of being selected and every possible sample of that size is equally likely

Question 502

[1 marks]sampling
State one method by which a random sample can be obtained from a population.

Answer this when you sit the paper.

Question 503

[3 marks]sampling
A school was asked to send 10 students on an exchange programme and to supply their names within 3 days. The head chose 10 students from among those who already had valid passports. Which sampling method did the head use?
  1. ASimple random sampling, since no student was named by the head before the selection was made
  2. BConvenience sampling, since the students were taken from the group that was quickest and easiest to reach
  3. CStratified sampling, since the school roll was divided into groups first and students were then drawn from each group in proportion to its size
  4. DSystematic sampling, since the students were taken at a fixed interval down an ordered list of names

Question 504

[3 marks]sampling
A head of school had to name 10 students for an exchange programme within 3 days, and chose 10 from among those students who already had valid passports. Would this give a random sample?
  1. ANo, because students without a valid passport had no chance at all of being chosen
  2. BYes, because the head did not know in advance which of the passport holders would be picked
  3. CYes, because 10 students is a large enough number for the selection to even out any bias
  4. DNo, because a sample of 10 students is too small for any selection method to count as random

Question 601

[2 marks]Poisson distribution
An insurance company receives claims at random at an average rate of 3 a week. Which feature of the claims makes a Poisson model appropriate here?
  1. AThe claims are received singly, at random and independently, at a constant average rate
  2. BThe claims are received in a fixed number of independent trials, each with the same probability
  3. CThe size of each claim is normally distributed about its own mean value
  4. DThe number of claims received can never exceed the average of 3 a week by very much

Question 602

[3 marks]Poisson distribution
An insurance company receives on average 3 claims in a week, the claims occurring at random and independently. Find the probability that the company receives at least 2 claims in a given week, to 3 decimal places.

Answer this when you sit the paper.

Question 603

[2 marks]Poisson distribution
An insurance company receives on average 3 claims a week and works a 5 day week. What Poisson mean should be used for the number of claims received in one working day?
  1. A15, since the weekly mean of 3 applies to each of the 5 working days
  2. B1.7, since the mean for a day is the square root of the mean for a week
  3. C3, since the average rate of claims does not change from one period to another
  4. D0.6, since the weekly mean of 3 is shared over the 5 working days

Question 604

[3 marks]Poisson distribution
An insurance company receives on average 3 claims a week and works 5 days in a week, so claims arrive at a mean rate of 0.6 a day. Find the probability that the company receives exactly one claim in a day, to 3 decimal places.

Answer this when you sit the paper.

Question 605

[3 marks]Poisson distribution
An insurance company receives claims at random at an average rate of 3 a week. Find the probability that it receives a total of exactly 2 claims during 3 consecutive weeks, to 3 decimal places.

Answer this when you sit the paper.

Question 606

[3 marks]Poisson distribution
For an insurance company receiving claims at a mean rate of 3 a week, the probability of at least 2 claims in a given week is 0.8009, and the weeks are independent. Find the probability that at least 2 claims are received in exactly one of 3 consecutive weeks, to 3 decimal places.

Answer this when you sit the paper.

Question 701

[2 marks]regression and correlation
For 10 pairs of readings of speed XX and fuel used YY, ∑x=955\sum x = 955, ∑x2=97 025\sum x^2 = 97\,025 and n=10n = 10. Find SxxS_{xx}.

Answer this when you sit the paper.

Question 702

[2 marks]regression and correlation
For 10 pairs of readings of speed XX and fuel used YY, ∑x=955\sum x = 955, ∑y=106\sum y = 106, ∑xy=10 810\sum xy = 10\,810 and n=10n = 10. Find SxyS_{xy}.

Answer this when you sit the paper.

Question 703

[3 marks]regression and correlation
For 10 pairs of readings of speed XX km/hr and fuel used YY litres, xˉ=95.5\bar x = 95.5, yˉ=10.6\bar y = 10.6, Sxx=5 822.5S_{xx} = 5\,822.5 and Sxy=687S_{xy} = 687. Find the equation of the regression line of YY on XX, giving the coefficients to 3 decimal places.

Answer this when you sit the paper.

Question 704

[3 marks]regression and correlation
The regression line of fuel used YY litres on speed XX km/hr, fitted to speeds between 60 and 140 km/hr, is Y=−0.668+0.118XY = -0.668 + 0.118X. Estimate the amount of fuel likely to be used when travelling at 105 km/hr, in litres to 1 decimal place.

Answer this when you sit the paper.

Question 705

[3 marks]regression and correlation
A regression line of fuel used on speed was fitted to ten readings taken at speeds between 60 and 140 km/hr. Why can the line not be relied on to estimate the fuel used at 50 km/hr?
  1. ABecause the regression line of YY on XX can only be used for values of XX that lie above the mean speed of the ten readings
  2. BBecause the line gives a negative estimate at 50 km/hr, which cannot be a quantity of fuel
  3. CBecause the product moment correlation coefficient is too far from 1 for any estimate to be made
  4. DBecause 50 km/hr lies outside the range of the speeds observed, so using the line there is extrapolation

Question 706

[3 marks]regression and correlation
For 10 pairs of readings of speed and fuel used, Sxx=5 822.5S_{xx} = 5\,822.5, Syy=88.4S_{yy} = 88.4 and Sxy=687S_{xy} = 687. Find the product moment correlation coefficient, to 3 decimal places.

Answer this when you sit the paper.

Question 801

[3 marks]confidence intervals
What is meant by a 95 % confidence interval for a population mean?
  1. AAn interval whose width is fixed at 95 % of the standard deviation of the sample that produced it
  2. BAn interval containing 95 % of the individual values in the population from which the sample was drawn, centred on the observed sample mean
  3. CAn interval, calculated from a sample, constructed by a method that captures the population mean in 95 % of repeated samples
  4. DAn interval within which 95 % of all possible sample means would fall if the population mean were zero

Question 802

[2 marks]confidence intervals
A random sample of 50 buses gave a mean of 70 passengers with a standard deviation of 4. Find the standard error of the sample mean, to 4 decimal places.

Answer this when you sit the paper.

Question 803

[3 marks]confidence intervals
A random sample of 50 buses gave a mean of 70 passengers with a standard deviation of 4, so the standard error of the mean is 0.5657. Calculate the 95 % confidence interval for the mean number of passengers in a bus, giving both limits to 1 decimal place.

Answer this when you sit the paper.

Question 804

[3 marks]confidence intervals
The number of passengers in a bus is normally distributed with mean 70 and standard deviation 4. Calculate the probability that a randomly chosen bus had fewer than 65 passengers, to 3 decimal places.

Answer this when you sit the paper.

Question 805

[2 marks]confidence intervals
In constructing a symmetric confidence interval for a mean at the 90 % level, state the value of zz that should be used, to 3 decimal places.

Answer this when you sit the paper.

Question 806

[3 marks]confidence intervals
The number of passengers in a bus has standard deviation 4. Calculate the sample size nn that should be taken so that one is 90 % confident that the sample mean will be within 0.8 of the true mean, using z=1.645z = 1.645.

Answer this when you sit the paper.

Question 901

[3 marks]hypothesis testing and chi-squared
What is the difference between a 1 tailed test and a 2 tailed test of a parameter?
  1. AA 1 tailed test is carried out at the 5 % level and a 2 tailed test is always carried out at the 10 % level
  2. BA 1 tailed test can only reject the null hypothesis, while a 2 tailed test can reject either hypothesis
  3. CA 1 tailed test has an alternative naming a direction, so the whole significance level sits in one tail; a 2 tailed test splits it between both tails
  4. DA 1 tailed test uses a sample from a single population, while a 2 tailed test compares samples drawn from two separate populations at the same significance level

Question 902

[2 marks]hypothesis testing and chi-squared
In a survey of newspaper readership the row total for the Northern province is 150, the column total for the newspaper 'current' is 160, and the grand total is 560. Calculate the expected frequency for Northern province readers of 'current', to 2 decimal places.

Answer this when you sit the paper.

Question 903

[2 marks]hypothesis testing and chi-squared
In a survey of newspaper readership the row total for the Southern province is 220, the column total for the newspaper 'News' is 190, and the grand total is 560. Calculate the expected frequency for Southern province readers of 'News', to 2 decimal places.

Answer this when you sit the paper.

Question 904

[2 marks]hypothesis testing and chi-squared
A chi-squared test of association is carried out on a contingency table with 3 rows (three provinces) and 3 columns (three newspapers). State the number of degrees of freedom.

Answer this when you sit the paper.

Question 905

[3 marks]hypothesis testing and chi-squared

A survey of newspaper readership in 3 provinces gave the observed frequencies below, with a grand total of 560.

ProvincetodaycurrentNews
Northern556530
Central804862
Southern754798

Using expected frequencies of 56.25, 42.86, 50.89 (Northern), 71.25, 54.29, 64.46 (Central) and 82.50, 62.86, 74.64 (Southern), calculate the chi-squared test statistic, to 1 decimal place.

Answer this when you sit the paper.

Question 906

[2 marks]hypothesis testing and chi-squared
A chi-squared test of association between province and newspaper preference gives a test statistic of 33.9 on 4 degrees of freedom. The 5 % critical value of χ2\chi^2 with 4 degrees of freedom is 9.488. What is the conclusion at the 5 % level of significance?
  1. ANo conclusion can be drawn, because a chi-squared statistic above 30 lies outside the range of the tables
  2. BReject the null hypothesis: there is evidence of an association between province and newspaper preference
  3. CAccept the null hypothesis: the two variables are independent, since 33.9 is well above the critical value
  4. DReject the null hypothesis: the test proves that living in a province causes a reader to choose one newspaper

Question 907

[2 marks]hypothesis testing and chi-squared
For a chi-squared test of association carried out at the 5 % level of significance on 4 degrees of freedom, state the critical value.

Answer this when you sit the paper.

Question 1001

[1 marks]normal distribution
The mass of a randomly chosen key follows a normal distribution with mean 12 g and variance 9. State the standard deviation of the mass of a key, in grams.

Answer this when you sit the paper.

Question 1002

[2 marks]normal distribution
The mass of a key-holder is N(20,16)N(20, 16) and the mass of a key is N(12,9)N(12, 9), all items independent. State the mean and the variance of the combined mass of 2 randomly chosen key-holders and 3 randomly chosen keys.

Answer this when you sit the paper.

Question 1003

[3 marks]normal distribution
The combined mass of 2 randomly chosen key-holders and 3 randomly chosen keys is normally distributed with mean 76 g and variance 59. Find the probability that this combined mass is greater than 78 g, to 3 decimal places.

Answer this when you sit the paper.

Question 1004

[2 marks]normal distribution
The mass of a key-holder is N(20,16)N(20, 16) and the mass of a key is N(12,9)N(12, 9), all items independent. Let DD be the combined mass of 3 key-holders minus the combined mass of 6 keys. State the mean and the variance of DD.

Answer this when you sit the paper.

Question 1005

[3 marks]normal distribution
For independent normal masses, DD = (mass of 3 key-holders) - (mass of 6 keys) has mean -12 and variance 102. Find the probability that the combined mass of 3 key-holders is greater than the combined mass of 6 keys, to 3 decimal places.

Answer this when you sit the paper.

Question 1006

[2 marks]normal distribution
The mass MM of a key-holder is N(20,16)N(20, 16) and the mass mm of a key is N(12,9)N(12, 9), independently. What is the variance of M−2mM - 2m?
  1. A16 - 18 = -2, because subtracting a quantity subtracts its variance as well as its mean
  2. B16 + 18 = 34, because the variance of the key is doubled along with its mean
  3. C16 + 36 = 52, because the multiplier 2 multiplies the key's variance by 222^2
  4. D16 + 9 = 25, because the multiplier changes the mean of the key but not its variance

Question 1007

[3 marks]normal distribution
For independent normal masses, D=M−2mD = M - 2m where MM is N(20,16)N(20, 16) and mm is N(12,9)N(12, 9), so DD has mean -4 and variance 52. Find the probability that a randomly chosen key-holder is more than twice the mass of a randomly chosen key, to 3 decimal places.

Answer this when you sit the paper.

Question 1101

[2 marks]grouped data and measures of location
The amounts 76 motorists spent on petrol are grouped in the classes 0≤x<500 \leq x < 50, 50≤x<10050 \leq x < 100, 100≤x<150100 \leq x < 150, 150≤x<200150 \leq x < 200, 200≤x<250200 \leq x < 250 and 250≤x<300250 \leq x < 300 dollars. State the mid-points of the six classes.

Answer this when you sit the paper.

Question 1102

[2 marks]grouped data and measures of location

76 motorists recorded the amount they spent on petrol in a month.

petrol purchases ($)number of motorists
0≤x<500 \leq x < 504
50≤x<10050 \leq x < 10011
100≤x<150100 \leq x < 1508
150≤x<200150 \leq x < 20016
200≤x<250200 \leq x < 25022
250≤x<300250 \leq x < 30015

Using the class mid-points 25, 75, 125, 175, 225 and 275, calculate the mean amount spent, in dollars to 2 decimal places.

Answer this when you sit the paper.

Question 1103

[2 marks]grouped data and measures of location

76 motorists recorded the amount they spent on petrol in a month.

petrol purchases ($)number of motorists
0≤x<500 \leq x < 504
50≤x<10050 \leq x < 10011
100≤x<150100 \leq x < 1508
150≤x<200150 \leq x < 20016
200≤x<250200 \leq x < 25022
250≤x<300250 \leq x < 30015

State the number of motorists who spent less than $200.

Answer this when you sit the paper.

Question 1104

[3 marks]grouped data and measures of location

76 motorists recorded the amount they spent on petrol in a month.

petrol purchases ($)number of motorists
0≤x<500 \leq x < 504
50≤x<10050 \leq x < 10011
100≤x<150100 \leq x < 1508
150≤x<200150 \leq x < 20016
200≤x<250200 \leq x < 25022
250≤x<300250 \leq x < 30015

Calculate the median amount spent, in dollars to 2 decimal places.

Answer this when you sit the paper.

Question 1105

[3 marks]grouped data and measures of location
For the petrol spending of 76 motorists, the class mid-points and frequencies give ∑fx=13 800\sum fx = 13\,800 and ∑fx2=2 927 500\sum fx^2 = 2\,927\,500, with mean 181.58181.58. Calculate the standard deviation, in dollars to 2 decimal places.

Answer this when you sit the paper.

Question 1106

[2 marks]grouped data and measures of location
A histogram is drawn for petrol spending grouped in six classes that are all 50 dollars wide, with frequencies 4, 11, 8, 16, 22 and 15. What should the height of each bar represent?
  1. AThe frequency density, which here is simply proportional to the frequency because every class has the same width
  2. BThe class mid-point, so that the tallest bar stands over the largest amount of money spent
  3. CThe frequency divided by the total of 76, so that the heights of all six bars add up to 1
  4. DThe cumulative frequency up to the upper boundary of that class, so that the bars always rise steadily from left to right across the display

Question 1107

[2 marks]grouped data and measures of location
In a histogram of petrol spending the modal class is 200≤x<250200 \leq x < 250 with frequency 22; the class below it has frequency 16 and the class above it has frequency 15, and every class is 50 wide. Estimate the mode, in dollars to 2 decimal places.

Answer this when you sit the paper.

Question 1201

[2 marks]time series
What is meant by seasonal variation in a time series?
  1. AThe gradual long term movement of the series in one direction, which is what is left once the short term swings have been removed
  2. BThe unpredictable one-off movement caused by an event outside the pattern, such as a strike
  3. CThe difference between the largest and the smallest value the series takes over the whole period
  4. DThe regular movement that repeats at the same point of every cycle, quarter by quarter or month by month

Question 1202

[2 marks]time series
What is meant by the trend in time series analysis?
  1. AThe pattern that repeats itself in the same quarter of each successive year of the series
  2. BThe underlying long term direction of the series once the seasonal and irregular movements are smoothed away
  3. CThe largest of the seasonal variations in the series, which sets the direction the whole series must then take over time
  4. DThe average of the values within any one year of the series, plotted at the middle of that year

Question 1203

[2 marks]time series
A pharmacy handled 1 700, 3 450, 2 800 and 2 300 customers in the four quarters of 2007, and 2 100 in the first quarter of 2008. Calculate the first 4-point moving average of this series.

Answer this when you sit the paper.

Question 1204

[2 marks]time series

A pharmacy handled these numbers of customers by quarter.

yearquarternumber of customers
200711 700
200723 450
200732 800
200742 300
200812 100
200823 500
200832 000
200842 000
200912 600
200924 600
200933 850
200943 800

Calculate the last of the nine 4-point moving averages of this series.

Answer this when you sit the paper.

Question 1205

[3 marks]time series
The first two 4-point moving averages of a quarterly series are 2 562.5 and 2 662.5. Calculate the first centred moving average, and state the quarter it is plotted at, given that the series starts at the first quarter of 2007.

Answer this when you sit the paper.

Question 1206

[2 marks]time series
The last two 4-point moving averages of a quarterly pharmacy series are 3 262.5 and 3 712.5. Calculate the last centred moving average.

Answer this when you sit the paper.

Question 1207

[3 marks]time series
The centred moving averages of a pharmacy's quarterly customer numbers run 2 612.5, 2 668.75, 2 575, 2 437.5, 2 462.5, 2 662.5, 3 031.25 and 3 487.5, from the third quarter of 2007 to the second quarter of 2009. What do they say about the trend?
  1. AThe trend dips a little through 2008 and then rises steadily, so customer numbers are increasing overall
  2. BThe trend falls steadily throughout, so the pharmacy is handling fewer customers year on year
  3. CThe trend rises and falls with each quarter, which is the seasonal pattern in the data rather than a trend in the series
  4. DThe series has no trend at all, since the values both fall and rise over the period covered

More sittings of this paper

The answers, and why they are the answers

Sit the paper here to see which ones you got right. Danho explains every question, keeps your score, and works without a connection.