# Who knows something about statistics?

**URL:** <https://boards.straightdope.com/t/who-knows-something-about-statistics/592077>\
**Category:** Factual Questions\
**Created:** [August 9, 2011, 1:46pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077 "2011-08-09T13:46:29Z")\
**Posts on this page:** 15\
**Page:** 2

<div class="post-metadata">

**Author:** ![Wendell\_Wagner](https://avatars.discourse-cdn.com/v4/letter/w/8491ac/32.png) [@Wendell\_Wagner](https://boards.straightdope.com/u/Wendell_Wagner)\
**Post date:** [August 10, 2011, 3:08am UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/21 "2011-08-10T03:08:02Z")

</div>

Would you please tell us what the study was trying to show? What was the hypothesis for this experiment?

---

<div class="post-metadata">

**Author:** ![mr.jp](https://avatars.discourse-cdn.com/v4/letter/m/48db29/32.png) [@mr.jp](https://boards.straightdope.com/u/mr.jp)\
**Post date:** [August 10, 2011, 2:28pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/22 "2011-08-10T14:28:34Z")

</div>

> [@D18](#):
>
> Since this is the context, they are telling the reader that there was a statistically significantly difference in ages between the boys and girls. That may or may not introduce a confound into the study (that is any claim made about a difference between boys and girls may actually be attributable to a difference between younger and older children).

I thought so at first too. But the p value shown is way too high for that.

---

<div class="post-metadata">

**Author:** ![Andy\_L](https://avatars.discourse-cdn.com/v4/letter/a/c67d28/32.png) [@Andy\_L](https://boards.straightdope.com/u/Andy_L)\
**Post date:** [August 10, 2011, 3:43pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/23 "2011-08-10T15:43:33Z")

</div>

If I’ve found the correct paper (and I think I have), the surrounding sentences say

“Mean age was 14.28 ± 1.78 years for the study group.  
Among the 87 members of the study group, 64 were girls  
(73.56%) and 23 were boys (26.44%). Mean age of girls  
was 13.56 ± 0.79, and of the boys was 14.73 ± 1.02  
(p = 0.032). There were significantly more migraine sufferers  
in the study group than in the contacted group  
(p\<0.001).”

I think the p value associated with the ages is the p-value for the null hypothesis that the boys and girls were selected from an underlying distribution of headache sufferers with a mean age of 14.28.

---

<div class="post-metadata">

**Author:** ![ultrafilter](https://avatars.discourse-cdn.com/v4/letter/u/3d9bf3/32.png) [@ultrafilter](https://boards.straightdope.com/u/ultrafilter)\
**Post date:** [August 10, 2011, 3:52pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/24 "2011-08-10T15:52:55Z")

</div>

How about a link to the paper?

---

<div class="post-metadata">

**Author:** ![Chessic\_Sense](https://avatars.discourse-cdn.com/v4/letter/c/7c8e57/32.png) [@Chessic\_Sense](https://boards.straightdope.com/u/Chessic_Sense)\
**Post date:** [August 10, 2011, 4:09pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/25 "2011-08-10T16:09:27Z")

</div>

> [@Smeghead](#):
>
> No. It’s not a measure of uncertainty of the measurements. It’s a measurement of probability. It’s saying that there’s a 3.2% chance that these _exact_ numbers could have come about through chance alone, even if there were no real effect.

But that’s only true if the age was hypothesized to be a dependent variable. If they just scooped up some middle schoolers and polled them on their ages, then it makes no sense to put a p value on it, because it’s independent.

Secondly, it’s not “these exact numbers”, but rather “these exact numbers or more” I’m sure the chances of getting those _exact_ numbers are infinitesimally small.

---

<div class="post-metadata">

**Author:** ![ultrafilter](https://avatars.discourse-cdn.com/v4/letter/u/3d9bf3/32.png) [@ultrafilter](https://boards.straightdope.com/u/ultrafilter)\
**Post date:** [August 10, 2011, 4:21pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/26 "2011-08-10T16:21:08Z")

</div>

> [@Chessic\_Sense](#):
>
> But that’s only true if the age was hypothesized to be a dependent variable. If they just scooped up some middle schoolers and polled them on their ages, then it makes no sense to put a p value on it, because it’s independent.

That’s simply not true. Any time you have two randomly chosen groups, you can do a test to compare whatever characteristics you’re interested in.

---

<div class="post-metadata">

**Author:** ![Buck\_Godot](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/buck_godot/32/6573_2.png) [@Buck\_Godot](https://boards.straightdope.com/u/Buck_Godot)\
**Post date:** [August 10, 2011, 8:37pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/27 "2011-08-10T20:37:29Z")

</div>

Here is the way I would interpret it.

The age of the Girls had mean 13.56 and standard deviation of 0.79  
The age of the Boys had mean 14.73 and standard deviation of 1.02

If there were no real systematic difference between the ages of the Boys and the Girls selected for the study, the probability that we would get such a large difference by chance is 3.2%

The numbers work out right if there were about 13 boys and 13 girls in the study.

As far as how to interpret it, assuming there is no issue of multiple comparisons (see XKCD link above), I like the following cutoffs

p\>0.1 Ignore result as probably just random chance  
0.01\<p\<0.1 Result worthy of interest and investigation, but not conclusive.  
0.001\<p\<0.01 Results most likely real (assuming test assumptions hold)  
p\<0.001 Results definitely real (assuming test assumptions hold)

So in this case, I would say, that there is some indication that there was a selection bias in the study for choosing older boys than girls, but it could just be chance.

---

<div class="post-metadata">

**Author:** ![Buck\_Godot](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/buck_godot/32/6573_2.png) [@Buck\_Godot](https://boards.straightdope.com/u/Buck_Godot)\
**Post date:** [August 10, 2011, 8:43pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/28 "2011-08-10T20:43:44Z")

</div>

> [@Chessic\_Sense](#):
>
> But that’s only true if the age was hypothesized to be a dependent variable. If they just scooped up some middle schoolers and polled them on their ages, then it makes no sense to put a p value on it, because it’s independent.

You still might want to do a test to make sure that your scooping of middle schoolers was unbiased. This result indicates that there is some evidence that it may not have been.

---

<div class="post-metadata">

**Author:** ![Andy\_L](https://avatars.discourse-cdn.com/v4/letter/a/c67d28/32.png) [@Andy\_L](https://boards.straightdope.com/u/Andy_L)\
**Post date:** [August 10, 2011, 9:33pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/29 "2011-08-10T21:33:05Z")

</div>

> [@ultrafilter](#):
>
> How about a link to the paper?

Sorry. I should have done that - here’s a link [http://www.springerlink.com/content/m05442w894377682/](http://www.springerlink.com/content/m05442w894377682/)

---

<div class="post-metadata">

**Author:** ![Buck\_Godot](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/buck_godot/32/6573_2.png) [@Buck\_Godot](https://boards.straightdope.com/u/Buck_Godot)\
**Post date:** [August 10, 2011, 10:11pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/30 "2011-08-10T22:11:29Z")

</div>

> [@Andy\_L](#):
>
> Sorry. I should have done that - here’s a link [http://www.springerlink.com/content/m05442w894377682/](http://www.springerlink.com/content/m05442w894377682/)

Looking at the paper in context, I’m pretty sure my initial reading of the results was right. They are using a different test than I expected and I made some mistakes in my calculations, so my sample size estimates were off otherwise what I said looks correct.

It is not uncommon for a study to report any anomalies in the sample selection that are found significant p\<0.05, even if it doesn’t have a great deal of effect on their conclusions.

---

<div class="post-metadata">

**Author:** ![ZenBeam](https://avatars.discourse-cdn.com/v4/letter/z/3ab097/32.png) [@ZenBeam](https://boards.straightdope.com/u/ZenBeam)\
**Post date:** [August 10, 2011, 10:20pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/31 "2011-08-10T22:20:39Z")

</div>

[QUOTE=Buck Godot]  
It is not uncommon for a study to report any anomalies in the sample selection that are found significant p\<0.05, even if it doesn’t have a great deal of effect on their conclusions.  
[/QUOTE]  
This makes the sentence in the OP make a lot more sense, at least to me.

Any ideas about “A total of 87 subjects completed the study: 64 girls (73.56%) and 23 boys (26.44%) (p = 0.016).” from the abstract in the link? Is that correct for that distribution of boys and girls when randomly selecting from a 50/50 mix? (Or the actual mix of 12 to 17 year-olds?)

---

<div class="post-metadata">

**Author:** ![Mean\_Mr.Mustard](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mean_mr.mustard/32/3865_2.png) [@Mean\_Mr.Mustard](https://boards.straightdope.com/u/Mean_Mr.Mustard)\
**Post date:** [August 11, 2011, 1:15am UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/32 "2011-08-11T01:15:54Z")

</div>

> [@ZenBeam](#):
>
> This makes the sentence in the OP make a lot more sense, at least to me.
> 
> Any ideas about “A total of 87 subjects completed the study: 64 girls (73.56%) and 23 boys (26.44%) (p = 0.016).” from the abstract in the link? Is that correct for that distribution of boys and girls when randomly selecting from a 50/50 mix? (Or the actual mix of 12 to 17 year-olds?)

I don’t think the subjects were randomly selected. I believe the participants were recruited based on their medical history and enrolled based on their willingness to participate.  
mmm

---

<div class="post-metadata">

**Author:** ![ZenBeam](https://avatars.discourse-cdn.com/v4/letter/z/3ab097/32.png) [@ZenBeam](https://boards.straightdope.com/u/ZenBeam)\
**Post date:** [August 11, 2011, 3:00am UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/33 "2011-08-11T03:00:47Z")

</div>

I’m not suggesting they were randomly selected, I’m just wondering if that’s the calculation the author’s made to get p=0.016. I’m also not suggesting that it makes sense to make that calculation, just wondering if they did.

---

<div class="post-metadata">

**Author:** ![mr.jp](https://avatars.discourse-cdn.com/v4/letter/m/48db29/32.png) [@mr.jp](https://boards.straightdope.com/u/mr.jp)\
**Post date:** [August 11, 2011, 3:39pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/34 "2011-08-11T15:39:10Z")

</div>

> [@Buck\_Godot](#):
>
> Here is the way I would interpret it.
> 
> The age of the Girls had mean 13.56 and standard deviation of 0.79  
> The age of the Boys had mean 14.73 and standard deviation of 1.02
> 
> If there were no real systematic difference between the ages of the Boys and the Girls selected for the study, the probability that we would get such a large difference by chance is 3.2%

I would also read it like that, but it doesn’t add up. I calculate the probability of getting such a large age difference by chance to be much lower.

---

<div class="post-metadata">

**Author:** ![Buck\_Godot](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/buck_godot/32/6573_2.png) [@Buck\_Godot](https://boards.straightdope.com/u/Buck_Godot)\
**Post date:** [August 11, 2011, 10:08pm UTC](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077/35 "2011-08-11T22:08:05Z")

</div>

> [@ZenBeam](#):
>
> This makes the sentence in the OP make a lot more sense, at least to me.
> 
> Any ideas about “A total of 87 subjects completed the study: 64 girls (73.56%) and 23 boys (26.44%) (p = 0.016).” from the abstract in the link? Is that correct for that distribution of boys and girls when randomly selecting from a 50/50 mix? (Or the actual mix of 12 to 17 year-olds?)

This is a bit more unclear. It can’t be assuming a 50/50 mix since that p-value would be much more extreme. I suspect that it might be the p-value for the difference in the percentages of those that enrolled in the study vs those that completed it, but it’s really not clear. They don’t indicate how many boys and girls entered the study so I can’t check.

> [@mr.jp](#):
>
> I would also read it like that, but it doesn’t add up. I calculate the probability of getting such a large age difference by chance to be much lower.

I agree that a t-test gives far too significant a p-value. But the methods section of the paper says they used a Wilcoxon rank test rather than a t-test to estimate the difference in ages, so we can’t recompute their results with just with the information provided. If there were a lot of ties in the ages (likely if ages were truncated by years) it’s possible that they could get that p-value.

[Previous page](https://boards.straightdope.com/t/who-knows-something-about-statistics/592077.md?page=1)
