题目
A French test and a Spanish test were sat by students. The table below shows their marks.
| Student | A | B | C | D | E | F | G | H | I | J | K |
|---|---|---|---|---|---|---|---|---|---|---|---|
| French mark | |||||||||||
| Spanish mark |
Greg says that if these points were plotted on a scatter diagram, then the point would be an outlier because is an outlier for the Spanish marks.
An outlier is defined as a value that is
or
(a) Show that is an outlier for the Spanish marks.
Ignoring the point , Greg calculated the following summary statistics.
(b) Use these summary statistics to show that the equation of the least squares regression line of on for the remaining students is
where the values of the intercept and gradient are given to significant figures. You must show your working.
(c) Give an interpretation of the gradient of the regression line.
Two further students sat the French test but missed the Spanish test.
(d) Using the equation given in part (b), estimate
(i) a Spanish mark for the student who scored marks in their French test,
(ii) a Spanish mark for the student who scored marks in their French test.
(e) State, giving a reason, which of the two estimates found in part (d) would be the more reliable estimate.
解答
(a)
解法一
思路
展开
把 Spanish marks 排序,找上下四分位数,再算上 outlier boundary。只要 超过上界,就说明它是 outlier。
答题过程
展开
For the Spanish marks,
The upper outlier boundary is
Since
is an outlier for the Spanish marks.
(b)
解法一
思路
展开
回归线 中,
然后用均值点求 。
答题过程
展开
The gradient is
For the remaining students,
So
Therefore, to significant figures,
(c)
解法一
思路
展开
斜率 表示 French mark 每增加 分,预测 Spanish mark 平均增加 分。
答题过程
展开
For each extra mark in the French test, the Spanish test mark is predicted to increase by about marks.
(d)(i)
解法一
思路
展开
把 代入回归方程。
答题过程
展开
When ,
The estimate is
(d)(ii)
解法一
思路
展开
把 代入同一个回归方程。
答题过程
展开
When ,
The estimate is
(e)
解法一
思路
展开
剩余 个学生的 French marks 范围是 到 。 在范围内,是 interpolation; 在范围外,是 extrapolation。因此 对应的估计更可靠。
答题过程
展开
The estimate for the student who scored in the French test is more reliable.
This is because is within the range of the French marks used to form the regression line, whereas is outside this range.