Danho
ZIMSEC A Level · 9164/4 · J2013

Statistics Paper 4 June 2013

Questions
43
Total marks
96
Syllabus code
9164/4

Sit this paper online

Questions
43
Pass mark
26
Sit this paper

Answer every question in the printed order, get marked at the end, then see the answers.

The questions

Section a

Section a, Question 2

[2 marks]Histograms and grouped data

The table shows the number of children below the age of 15 known to have suffered from measles in 2009 in a certain village.

Age (in years)Number of reported cases
Under 114
1 - 233
3 - 435
5 - 939
10 - 145

Age is a continuous quantity, so the class boundaries are 0 to 1, 1 to 2.5, 2.5 to 4.5, 4.5 to 9.5 and 9.5 to 14.5. Calculate the mean age, in years, of the children who suffered from the disease.

Answer this when you sit the paper.

[2 marks]Histograms and grouped data

The table shows the number of children below the age of 15 known to have suffered from measles in 2009 in a certain village.

Age (in years)Number of reported cases
Under 114
1 - 233
3 - 435
5 - 939
10 - 145

In a histogram of these data the bars are drawn to frequency density, since the classes are of unequal width. Taking the boundaries of the 5 - 9 class as 4.5 to 9.5, find the frequency density of that class.

Answer this when you sit the paper.

[3 marks]Tree diagrams and conditional probability

During the 2010 World Cup in a certain city, the probability that there was electricity on any particular day was 13\dfrac{1}{3}. In the case that there was no electricity, a generator would be switched on. Independently, the probability that John watched a soccer match being screened live was 14\dfrac{1}{4}.

Given that John watched a soccer match, find the probability that there was no electricity. Give your answer as a fraction in its lowest terms.

Answer this when you sit the paper.

[1 marks]Tree diagrams and conditional probability

During the 2010 World Cup in a certain city, the probability that there was electricity on any particular day was 13\dfrac{1}{3}. In the case that there was no electricity, a generator would be switched on. Independently, the probability that John watched a soccer match being screened live was 14\dfrac{1}{4}.

Find the probability that on a particular day there was no electricity and John watched a soccer match.

  1. A14\dfrac{1}{4}
  2. B16\dfrac{1}{6}
  3. C23\dfrac{2}{3}
  4. D112\dfrac{1}{12}
[1 marks]Estimation and confidence intervals

An ice-cream vendor records his daily takings (xx) over a period of 30 days. The results are summarised by ∑x=900\sum x = 900 and ∑x2=34 000\sum x^{2} = 34\ 000.

Find the unbiased estimate of the population mean.

Answer this when you sit the paper.

[2 marks]Estimation and confidence intervals

An ice-cream vendor records his daily takings (xx) over a period of 30 days. The results are summarised by ∑x=900\sum x = 900 and ∑x2=34 000\sum x^{2} = 34\ 000.

Find the unbiased estimate of the population variance.

Answer this when you sit the paper.

[3 marks]Estimation and confidence intervals

An ice-cream vendor records his daily takings (xx) over a period of 30 days. The results are summarised by ∑x=900\sum x = 900 and ∑x2=34 000\sum x^{2} = 34\ 000.

The unbiased estimate of the population variance is 241.38. Assuming the daily takings are normally distributed and using z=1.96z = 1.96, calculate the upper limit of the 95% confidence interval for the mean amount he receives.

Answer this when you sit the paper.

Section a, Question 3

[1 marks]Geometric distribution

The probability that a learner driver passes the driving test at the Vehicle Inspection Department is 14\dfrac{1}{4}, and the attempts are independent of one another. A learner counts the number of attempts made up to and including the one on which the test is passed.

State the statistical distribution which best models the number of attempts.

  1. AGeometric, because the learner keeps attempting until the first pass and every attempt carries the same chance of succeeding
  2. BBinomial, because each attempt ends in either a pass or a fail and the number of attempts is settled before the learner starts out
  3. CPoisson, because a pass is a rare event that happens at a steady average rate over the run of attempts made
  4. DNormal, because the number of attempts is a measurement spread symmetrically about its own mean value
[2 marks]Geometric distribution

The probability that a learner driver passes the driving test at the Vehicle Inspection Department is 14\dfrac{1}{4}, and the attempts are independent of one another. A learner counts the number of attempts made up to and including the one on which the test is passed.

Find the mean number of attempts.

Answer this when you sit the paper.

[2 marks]Geometric distribution

The probability that a learner driver passes the driving test at the Vehicle Inspection Department is 14\dfrac{1}{4}, and the attempts are independent of one another. A learner counts the number of attempts made up to and including the one on which the test is passed.

Find the variance of the number of attempts.

Answer this when you sit the paper.

[3 marks]Geometric distribution

The probability that a learner driver passes the driving test at the Vehicle Inspection Department is 14\dfrac{1}{4}, and the attempts are independent of one another. A learner counts the number of attempts made up to and including the one on which the test is passed.

Find the smallest value of nn for which the probability that the learner needs only nn or fewer attempts is at least 0.7.

Answer this when you sit the paper.

[2 marks]Continuous random variables

A continuous random variable X has probability density function

f(x)={120≤x<0.515(3−x),0.5≤x≤30,otherwise.f(x)=\begin{cases}\dfrac{1}{2} & 0 \le x < 0.5\\[4pt] \dfrac{1}{5}(3-x), & 0.5 \le x \le 3\\[4pt] 0, & \text{otherwise.}\end{cases}

Which of these describes the graph of f(x)f(x)?

  1. AA horizontal segment at height 0.5 from x=0x=0 to x=0.5x=0.5, then a straight line falling from (0.5, 0.5)(0.5,\ 0.5) to (3, 0)(3,\ 0), and zero elsewhere
  2. BA horizontal segment at height 0.5 from x=0x=0 to x=0.5x=0.5, then a straight line rising from (0.5, 0.5)(0.5,\ 0.5) to (3, 1)(3,\ 1), and zero elsewhere
  3. CA horizontal segment at height 0.2 from x=0x=0 to x=0.5x=0.5, then a straight line falling from (0.5, 0.2)(0.5,\ 0.2) to (3, 0)(3,\ 0), and zero elsewhere
  4. DA horizontal segment at height 0.5 from x=0x=0 to x=0.5x=0.5, then a curve falling steeply from (0.5, 0.5)(0.5,\ 0.5) towards (3, 0)(3,\ 0), and zero elsewhere
[3 marks]Continuous random variables

A continuous random variable X has probability density function

f(x)={120≤x<0.515(3−x),0.5≤x≤30,otherwise.f(x)=\begin{cases}\dfrac{1}{2} & 0 \le x < 0.5\\[4pt] \dfrac{1}{5}(3-x), & 0.5 \le x \le 3\\[4pt] 0, & \text{otherwise.}\end{cases}

Find the median of X.

Answer this when you sit the paper.

[3 marks]Continuous random variables

A continuous random variable X has probability density function

f(x)={120≤x<0.515(3−x),0.5≤x≤30,otherwise.f(x)=\begin{cases}\dfrac{1}{2} & 0 \le x < 0.5\\[4pt] \dfrac{1}{5}(3-x), & 0.5 \le x \le 3\\[4pt] 0, & \text{otherwise.}\end{cases}

Evaluate P(x<1.2)P(x < 1.2).

Answer this when you sit the paper.

[2 marks]Hypothesis testing of a mean

A machine on a production line is set to make components with a mean diameter of 2 cm. A random sample of 10 components had their diameters, in cm, measured:

2.17 1.93 2.02 1.97 2.00 2.01 2.02 1.89 1.99 2.01

Calculate the mean diameter of the sample, in cm.

Answer this when you sit the paper.

[3 marks]Hypothesis testing of a mean

A machine on a production line is set to make components with a mean diameter of 2 cm. A random sample of 10 components had their diameters, in cm, measured:

2.17 1.93 2.02 1.97 2.00 2.01 2.02 1.89 1.99 2.01

The ten readings give ∑x=20.01\sum x = 20.01 and ∑x2=40.0879\sum x^{2} = 40.0879. Calculate the unbiased estimate of the population standard deviation, in cm.

Answer this when you sit the paper.

[2 marks]Hypothesis testing of a mean

A machine on a production line is set to make components with a mean diameter of 2 cm. A random sample of 10 components had their diameters, in cm, measured:

2.17 1.93 2.02 1.97 2.00 2.01 2.02 1.89 1.99 2.01

The mean diameter is to be tested against the set value of 2 cm at the 5% level of significance, with the spread estimated from the sample itself. State the positive critical value of the test statistic.

Answer this when you sit the paper.

[2 marks]Hypothesis testing of a mean

A machine on a production line is set to make components with a mean diameter of 2 cm. A random sample of 10 components had their diameters, in cm, measured:

2.17 1.93 2.02 1.97 2.00 2.01 2.02 1.89 1.99 2.01

For these data the test statistic works out at 0.043 and the two-tailed 5% critical values on 9 degrees of freedom are ±2.262\pm 2.262. What is the conclusion of the test?

  1. AThe test statistic is larger than the critical value, so the null hypothesis is rejected and the mean diameter is taken to differ from the set 2 cm
  2. BThe test statistic is smaller than the critical value, so the null hypothesis is retained and the mean diameter is taken to be the set 2 cm
  3. CThe test statistic is positive rather than negative, so the null hypothesis is rejected and the mean diameter is taken to be above the set 2 cm
  4. DThe sample holds fewer than thirty readings, so the machine setting cannot be judged from a sample of this size and no conclusion follows

Section a, Question 4

[2 marks]Hypothesis testing of a proportion
Distinguish between a 1-tailed test and a 2-tailed test of a hypothesis.
  1. AA 1-tailed test puts half of the significance level into each tail of the distribution, while a 2-tailed test puts the whole of the significance level into one tail of it
  2. BA 1-tailed test states its alternative in one direction, so its rejection region lies wholly in one tail, while a 2-tailed test states a difference either way and splits the region
  3. CA 1-tailed test compares the means of two separate populations against each other, while a 2-tailed test compares one sample mean against a single stated population value
  4. DA 1-tailed test needs a larger sample than a 2-tailed test does, and for that reason its rejection region is placed wholly in the upper tail of the distribution whichever direction the alternative names
[2 marks]Hypothesis testing of a proportion

A political party claims that it commands 60% of the voters. To test this, a random sample of 300 potential voters were asked which party they would vote for, and 160 confirmed that they would vote for that party. The claim is tested at the 10% level of significance.

Calculate the sample proportion of voters who confirmed that they would vote for the party.

Answer this when you sit the paper.

[3 marks]Hypothesis testing of a proportion

A political party claims that it commands 60% of the voters. To test this, a random sample of 300 potential voters were asked which party they would vote for, and 160 confirmed that they would vote for that party. The claim is tested at the 10% level of significance.

Taking H0:p=0.6H_{0}: p = 0.6 against H1:p<0.6H_{1}: p < 0.6, and the sample proportion as 160300\dfrac{160}{300}, calculate the value of the test statistic z.

Answer this when you sit the paper.

[2 marks]Hypothesis testing of a proportion

A political party claims that it commands 60% of the voters. To test this, a random sample of 300 potential voters were asked which party they would vote for, and 160 confirmed that they would vote for that party. The claim is tested at the 10% level of significance.

The test statistic works out at z=−2.357z = -2.357 and the 10% one-tailed critical value is −1.282-1.282. What is the conclusion?

  1. AThe test statistic is below the critical value, so the null hypothesis is rejected: there is evidence that the party commands less than 60% of the voters
  2. BThe test statistic is above the critical value, so the null hypothesis is retained: the sample supports the party's claim to command 60% of the voters
  3. CThe test statistic is below the critical value, so the null hypothesis is retained: the sample gives no reason at all to doubt the claim made by the party
  4. DThe test statistic is below the 1% critical value of −2.326-2.326, so the claim is retained at the 10% level although it would be rejected at the 1% level
[1 marks]Poisson distribution

Mangoes are boxed into cartons, each containing 500 mangoes. The probability that any one mango is rotten is 0.002, independently of the others. A buyer returns any carton that contains 4 or more rotten mangoes.

Find the expected number of rotten mangoes per carton.

Answer this when you sit the paper.

[2 marks]Poisson distribution

Mangoes are boxed into cartons, each containing 500 mangoes. The probability that any one mango is rotten is 0.002, independently of the others. A buyer returns any carton that contains 4 or more rotten mangoes.

State the most appropriate statistical distribution for modelling the number of rotten mangoes in a carton, with the reason for it.

  1. AA normal distribution with mean 1 and variance 1, because 500 mangoes in a carton are more than enough for a continuous model to be used
  2. BA geometric distribution with p=0.002p = 0.002, because the count of interest is how many mangoes are inspected before the first rotten one turns up
  3. CA binomial distribution with n=4n = 4 and p=0.002p = 0.002, because the buyer returns the carton as soon as a fourth rotten mango is found in it
  4. DA Poisson distribution with mean 1, because the number of mangoes is large, the chance any one is rotten is small, and npnp is well below 5
[1 marks]Poisson distribution

Mangoes are boxed into cartons, each containing 500 mangoes. The probability that any one mango is rotten is 0.002, independently of the others. A buyer returns any carton that contains 4 or more rotten mangoes.

Two cartons of mangoes are chosen at random. Find the expected number of rotten mangoes in the two cartons together.

Answer this when you sit the paper.

[3 marks]Poisson distribution

Mangoes are boxed into cartons, each containing 500 mangoes. The probability that any one mango is rotten is 0.002, independently of the others. A buyer returns any carton that contains 4 or more rotten mangoes.

Find the probability that a carton of mangoes is not returned.

Answer this when you sit the paper.

[3 marks]Poisson distribution

Mangoes are boxed into cartons, each containing 500 mangoes. The probability that any one mango is rotten is 0.002, independently of the others. A buyer returns any carton that contains 4 or more rotten mangoes.

Two cartons of mangoes are chosen at random. Find the probability that between them they contain at least three rotten mangoes.

Answer this when you sit the paper.

[2 marks]Chi-squared test of association

The table shows the ownership of satellite dishes by different social classes in a randomly chosen sample of 150 households in a town.

social classown a satellite dishdo not own a satellite dish
Executive1510
Managerial238
Working5440

Assuming there is no association between ownership and social class, calculate the expected number of executive households that own a satellite dish.

Answer this when you sit the paper.

[2 marks]Chi-squared test of association

The table shows the ownership of satellite dishes by different social classes in a randomly chosen sample of 150 households in a town.

social classown a satellite dishdo not own a satellite dish
Executive1510
Managerial238
Working5440

State the number of degrees of freedom for a chi-squared test of association on this table.

Answer this when you sit the paper.

[3 marks]Chi-squared test of association

The table shows the ownership of satellite dishes by different social classes in a randomly chosen sample of 150 households in a town.

social classown a satellite dishdo not own a satellite dish
Executive1510
Managerial238
Working5440

The expected frequencies are 15.3 and 9.7 for Executive, 19.0 and 11.98 for Managerial, and 57.66 and 36.3 for Working. Calculate the value of the chi-squared test statistic.

Answer this when you sit the paper.

[2 marks]Chi-squared test of association

The table shows the ownership of satellite dishes by different social classes in a randomly chosen sample of 150 households in a town.

social classown a satellite dishdo not own a satellite dish
Executive1510
Managerial238
Working5440

The test is carried out at the 5% level of significance on 2 degrees of freedom. State the critical value of the chi-squared statistic.

Answer this when you sit the paper.

[2 marks]Chi-squared test of association

The table shows the ownership of satellite dishes by different social classes in a randomly chosen sample of 150 households in a town.

social classown a satellite dishdo not own a satellite dish
Executive1510
Managerial238
Working5440

The test statistic works out at 2.79 and the 5% critical value on 2 degrees of freedom is 5.991. What is the conclusion of the test?

  1. AThe calculated value is less than 5.991, so the null hypothesis is rejected and ownership of a dish is shown to be independent of social class in this town
  2. BThe calculated value is less than 5.991 but rests on 5 degrees of freedom, so nothing follows until a larger sample of households has been surveyed
  3. CThe calculated value is greater than 5.991, so the null hypothesis is rejected and there is evidence of an association between ownership and social class
  4. DThe calculated value is less than 5.991, so the null hypothesis is retained and there is no evidence of an association between ownership and social class

Section a, Question 5

[3 marks]Normal distribution

Oranges are sold in small pockets whose weights are normally distributed with mean μ\mu kg and standard deviation σ\sigma kg. The probability that a randomly chosen pocket weighs more than 3.2 kg is 0.1, and the probability that a randomly chosen pocket weighs less than 2.4 kg is 0.2.

Find σ\sigma, in kg.

Answer this when you sit the paper.

[2 marks]Normal distribution

Oranges are sold in small pockets whose weights are normally distributed with mean μ\mu kg and standard deviation σ\sigma kg. The probability that a randomly chosen pocket weighs more than 3.2 kg is 0.1, and the probability that a randomly chosen pocket weighs less than 2.4 kg is 0.2.

Find μ\mu, in kg.

Answer this when you sit the paper.

[2 marks]Normal distribution, sums and differences

Oranges are sold in small pockets whose weights are normally distributed with mean 2.7 kg and standard deviation 0.38 kg. Potatoes are sold in small pockets whose weights are normally distributed with mean 5 kg and standard deviation 0.6 kg. All the pockets are chosen independently of one another.

Find the mean of the combined weight of two orange pockets and one potato pocket, in kg.

Answer this when you sit the paper.

[3 marks]Normal distribution, sums and differences

Oranges are sold in small pockets whose weights are normally distributed with mean 2.7 kg and standard deviation 0.38 kg. Potatoes are sold in small pockets whose weights are normally distributed with mean 5 kg and standard deviation 0.6 kg. All the pockets are chosen independently of one another.

Find the probability that two randomly chosen orange pockets together weigh less than one randomly chosen pocket of potatoes.

Answer this when you sit the paper.

[3 marks]Normal distribution, sums and differences

Oranges are sold in small pockets whose weights are normally distributed with mean 2.7 kg and standard deviation 0.38 kg. Potatoes are sold in small pockets whose weights are normally distributed with mean 5 kg and standard deviation 0.6 kg. All the pockets are chosen independently of one another.

Find the probability that two randomly chosen orange pockets and one randomly chosen pocket of potatoes weigh more than 11 kg altogether.

Answer this when you sit the paper.

[2 marks]Scatter diagrams and regression

The manager of a clothing shop surveyed ten women, recording each woman's age X in years and her annual expenditure on clothes Y in dollars.

age (years) X18213645235325373032
expenditure (dollars) Y330300180120310150250150245190

Calculate the mean annual expenditure on clothes, in dollars.

Answer this when you sit the paper.

[3 marks]Scatter diagrams and regression

For ten women, the age X in years and the annual expenditure on clothes Y in dollars gave n=10n = 10, ∑x=320\sum x = 320, ∑y=2225\sum y = 2225, ∑x2=11 342\sum x^{2} = 11\,342, ∑y2=545 425\sum y^{2} = 545\,425 and ∑xy=64 430\sum xy = 64\,430.

Calculate the gradient of the regression line of Y on X.

Answer this when you sit the paper.

[3 marks]Scatter diagrams and regression

For ten women, the age X in years and the annual expenditure on clothes Y in dollars gave n=10n = 10, ∑x=320\sum x = 320, ∑y=2225\sum y = 2225, ∑x2=11 342\sum x^{2} = 11\,342, ∑y2=545 425\sum y^{2} = 545\,425 and ∑xy=64 430\sum xy = 64\,430.

The gradient of the regression line of Y on X is −6.143-6.143, and xˉ=32\bar{x}=32, yˉ=222.5\bar{y}=222.5. Which is the equation of that line?

  1. AY=−6.143X−419.08Y = -6.143X - 419.08
  2. BY=6.143X+419.08Y = 6.143X + 419.08
  3. CY=−6.143X+419.08Y = -6.143X + 419.08
  4. DY=−0.134X+226.80Y = -0.134X + 226.80
[2 marks]Scatter diagrams and regression

For ten women, the age X in years and the annual expenditure on clothes Y in dollars gave n=10n = 10, ∑x=320\sum x = 320, ∑y=2225\sum y = 2225, ∑x2=11 342\sum x^{2} = 11\,342, ∑y2=545 425\sum y^{2} = 545\,425 and ∑xy=64 430\sum xy = 64\,430.

The regression line of Y on X is Y=−6.143X+419.08Y = -6.143X + 419.08. Use it to estimate the amount, in dollars, likely to be spent on clothes by a 40 year old woman.

Answer this when you sit the paper.

[3 marks]Correlation

For ten women, the age X in years and the annual expenditure on clothes Y in dollars gave n=10n = 10, ∑x=320\sum x = 320, ∑y=2225\sum y = 2225, ∑x2=11 342\sum x^{2} = 11\,342, ∑y2=545 425\sum y^{2} = 545\,425 and ∑xy=64 430\sum xy = 64\,430.

Calculate the product-moment correlation coefficient between age and annual expenditure.

Answer this when you sit the paper.

[1 marks]Correlation
For a sample of ten women, the product-moment correlation coefficient between a woman's age and her annual expenditure on clothes was found to be −0.909-0.909. Comment on this value.
  1. AThe value is close to zero, so there is almost no linear correlation between a woman's age and the amount she spent each year on clothes
  2. BThe value is negative, so it shows that in this sample growing a year older is what caused each woman to spend less on her clothes
  3. CThe value is close to −1-1, so there is a strong negative linear correlation: in this sample, the older the woman the less she spent on clothes
  4. DThe value is close to −1-1, so there is only a weak negative linear correlation: in this sample a woman's age said very little about what she spent

More sittings of this paper

The answers, and why they are the answers

Sit the paper here to see which ones you got right. Danho explains every question, keeps your score, and works without a connection.