# A quick stats computation

**URL:** <https://boards.straightdope.com/t/a-quick-stats-computation/472477>\
**Category:** Factual Questions\
**Created:** [November 13, 2008, 3:22am UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477 "2008-11-13T03:22:35Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chessic\_Sense](https://avatars.discourse-cdn.com/v4/letter/c/7c8e57/32.png) [@Chessic\_Sense](https://boards.straightdope.com/u/Chessic_Sense)\
**Post date:** [November 13, 2008, 3:22am UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/1 "2008-11-13T03:22:35Z")

</div>

Is a 52-59% difference significant between 2 groups of 600 people each? I forget how to figure that out, exactly. A detailed answer would be appreciated.

Again, 52% of group one (N=600) and 59% of group two (N=600).

Thanks

---

<div class="post-metadata">

**Author:** ![Chessic\_Sense](https://avatars.discourse-cdn.com/v4/letter/c/7c8e57/32.png) [@Chessic\_Sense](https://boards.straightdope.com/u/Chessic_Sense)\
**Post date:** [November 13, 2008, 5:27pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/2 "2008-11-13T17:27:23Z")

</div>

Seriously? NO ONE?!

---

<div class="post-metadata">

**Author:** ![Santo\_Rugger](https://avatars.discourse-cdn.com/v4/letter/s/e95f7d/32.png) [@Santo\_Rugger](https://boards.straightdope.com/u/Santo_Rugger)\
**Post date:** [November 13, 2008, 5:40pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/3 "2008-11-13T17:40:10Z")

</div>

For a binomial distribution, the standard deviation (sigma) is sqrt (p \* q \* n), where p is probability, q is 1 - probability, and n is the number of trials.

sqrt (.52 \* (1-.52) \* 600) = 12.23  
sqrt (.59 \* (1-.59) \* 600) = 12.047

This means that your answer within one standard deviation will probably vary by about 12 for each case (out of 600). Three standard deviations are off by 36, and is 95% confident. (59-52)/100 \* 600 = 42.

Since 42 is bigger than 36, we can say these results are statistically significant.

I think.

> **[Binomial distribution](https://en.wikipedia.org/wiki/Binomial_distribution)**
>
> In probability theory and statistics, the binomial distribution with parameters n and p is the discrete probability distribution of the number of successes in a sequence of n independent experiments, each asking a yes–no question, and each with its own Boolean-valued outcome: success (with probability p) or failure (with probability 
>   
>     
>       
> q
> =
> 1
> −
> p
>       
>     
> {\\displaystyle q=1-p}
>   
> ). A single success/failure experiment is also called a Ber...

---

<div class="post-metadata">

**Author:** ![cmosdes](https://avatars.discourse-cdn.com/v4/letter/c/a587f6/32.png) [@cmosdes](https://boards.straightdope.com/u/cmosdes)\
**Post date:** [November 13, 2008, 5:55pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/4 "2008-11-13T17:55:35Z")

</div>

According to my Practical Data Analysis workbook notes, in order to determine if the % population fraction from two samples is the same, you’d need a sample size of:

n1 = n2 = 16 \* (%) \* (1 - %)/(delta%)^2

Assuming your expected % is around 55.5%, then plugging these numbers into the above formula I get:

n1 = n2 = 16 \* (0.555) \* (1 - 0.555)/(0.07)^2 = 806.449

In other words, you’d need a sample size of 806.449 in order for these to be statistically the same. Since you have fewer than that, the error in your % will be greater and therefore they are likely the same.

But now I see Santo Rugger posted just the opposite. I’ll check around a little more.

---

<div class="post-metadata">

**Author:** ![Santo\_Rugger](https://avatars.discourse-cdn.com/v4/letter/s/e95f7d/32.png) [@Santo\_Rugger](https://boards.straightdope.com/u/Santo_Rugger)\
**Post date:** [November 13, 2008, 6:17pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/5 "2008-11-13T18:17:54Z")

</div>

I’m trying to work out the probability density function, but the numbers are way to huge for my calculator or MatLab to handle (600!).

---

<div class="post-metadata">

**Author:** ![ultrafilter](https://avatars.discourse-cdn.com/v4/letter/u/3d9bf3/32.png) [@ultrafilter](https://boards.straightdope.com/u/ultrafilter)\
**Post date:** [November 13, 2008, 6:52pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/6 "2008-11-13T18:52:32Z")

</div>

I don’t trust the reasoning behind either of the answers presented so far. There’s a calculator [here](http://www.answersresearch.com/proportions.php) that will tell you whether there’s a statistically significant difference at a specified confidence level, if that’s all you need. If you need to be able to perform the test yourself, I believe you want to look at [Welch’s t-test](http://en.wikipedia.org/wiki/Welch%27s_t_test).

For the record, the calculator above specifies that your proportions are different at 95% confidence.

---

<div class="post-metadata">

**Author:** ![Santo\_Rugger](https://avatars.discourse-cdn.com/v4/letter/s/e95f7d/32.png) [@Santo\_Rugger](https://boards.straightdope.com/u/Santo_Rugger)\
**Post date:** [November 13, 2008, 7:01pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/7 "2008-11-13T19:01:35Z")

</div>

> [@ultrafilter](#):
>
> I don’t trust the reasoning behind either of the answers presented so far. There’s a calculator [here](http://www.answersresearch.com/proportions.php) that will tell you whether there’s a statistically significant difference at a specified confidence level, if that’s all you need. If you need to be able to perform the test yourself, I believe you want to look at [Welch’s t-test](http://en.wikipedia.org/wiki/Welch%27s_t_test).
> 
> For the record, the calculator above specifies that your proportions are different at 95% confidence.

How is the t-test different than what I did? IIRC, the t test shows the probability of the answer being in the “tail” of the distribution curve. Three sigmas is 95%, or a 2.5% chance of being in the “tail” of the curve. Four sigmas is 99%, so using the method I used, 12\*4 \> 46, hence the results not being statistically significant at 99%.

I’m horrible at statistics, and welcome the flaw in my logic being pointed out.

---

<div class="post-metadata">

**Author:** ![muttrox](https://avatars.discourse-cdn.com/v4/letter/m/a8b319/32.png) [@muttrox](https://boards.straightdope.com/u/muttrox)\
**Post date:** [November 13, 2008, 7:36pm UTC](https://boards.straightdope.com/t/a-quick-stats-computation/472477/8 "2008-11-13T19:36:55Z")

</div>

> [@cmosdes](#):
>
> According to my Practical Data Analysis workbook notes, in order to determine if the % population fraction from two samples is the same, you’d need a sample size of:
> 
> n1 = n2 = 16 \* (%) \* (1 - %)/(delta%)^2
> 
> Assuming your expected % is around 55.5%, then plugging these numbers into the above formula I get:
> 
> n1 = n2 = 16 \* (0.555) \* (1 - 0.555)/(0.07)^2 = 806.449
> 
> In other words, you’d need a sample size of 806.449 in order for these to be statistically the same. Since you have fewer than that, the error in your % will be greater and therefore they are likely the same.
> 
> But now I see Santo Rugger posted just the opposite. I’ll check around a little more.

I follow Santo Rugger’s and Ultrafilters reasoning, this one makes no sense to me.
