# Statistics and sample-size question

**URL:** <https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882>\
**Category:** Factual Questions\
**Created:** [September 7, 2026, 6:08pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882 "2026-09-07T18:08:06Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![hideousidiot](https://avatars.discourse-cdn.com/v4/letter/h/858c86/32.png) [@hideousidiot](https://boards.straightdope.com/u/hideousidiot)\
**Post date:** [September 7, 2026, 6:08pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/1 "2026-09-07T18:08:06Z")

</div>

According to an AI I consulted, and my own sense, education-levels in this country correlate closely with voting patterns: people with doctorates tend to vote Democratic much more than people who never graduated high school, and the four points polled in between those two extremes (did graduate high school, went to college but did not graduate, college grads, went to grad school but did not take a degree) increasingly voted Democratic at each gradation of education.

If you wish to dispute the AI on this, feel free (I have no reason to think it’s wrong), but assuming it’s right, I wonder what would result if we were able to poll sixty levels of education rather than just six. There are about sixty levels that can be marked (and probably more): each of the K-12 grades can be divided into “started” and “finished” (and divided by semester as well) as can college (which could be further subdivided by GPA, credits earned, etc.) . Grad school can be divided several different ways, each denoting a higher or lower level of completion, starting with the basic one of “got Masters degree” and “got doctorate.” And beyond doctorate, there are a few more levels, such as post-doctoral studies, multiple graduate degrees, etc.

My question is: assuming we could gather data at each of these levels of education, would they eventually show a smooth correlation between the sixty education-levels as there is between the six education-levels we now measure?

And if “yes” is the answer to that, how much data would be required at each level to eliminate glitches caused by small-sample size? I’m thinking we’d need more data at each level than the population actually holds. That is, are there enough people who’ve started fourth grade but did not finish it in the entire country to make for stable reliable polling results, even if you were able to poll each person who qualifies for incusion? I think not but am willing to be proven wrong.

---

<div class="post-metadata">

**Author:** ![OldGuy](https://avatars.discourse-cdn.com/v4/letter/o/3bc359/32.png) [@OldGuy](https://boards.straightdope.com/u/OldGuy)\
**Post date:** [September 7, 2026, 7:04pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/2 "2026-09-07T19:04:23Z")

</div>

It really depends on what you want. You’d need a lot more data to be confident that the relation was monotonic than say that a regression line was upward sloping. Also it depends on your model. If you could sample everyone – know how they voted and where their schooling ended, then it’s not a statistical question as you have the population data so their is no estimation whether the relation is monotonic; you’d now it (or not) for a fact.

---

<div class="post-metadata">

**Author:** ![Kent\_Clark](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kent_clark/32/105_2.png) [@Kent\_Clark](https://boards.straightdope.com/u/Kent_Clark)\
**Post date:** [September 7, 2026, 8:51pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/3 "2026-09-07T20:51:27Z")

</div>

Trying to mesaure the political views of a kindergarten student is pointless. And finding an adult with only a kindergarten education who votes is probably impossible. Pick a more appropriate age of politcal awareness and understanding for a baseline.

---

<div class="post-metadata">

**Author:** ![Francis\_Vaughan](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/francis_vaughan/32/3093_2.png) [@Francis\_Vaughan](https://boards.straightdope.com/u/Francis_Vaughan)\
**Post date:** [September 7, 2026, 9:34pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/4 "2026-09-07T21:34:32Z")

</div>

A very fine grained division may act to bring in confusing factors. If you add “started but not completed” into the mix there is a good chance that the reason someone didn’t finish becomes important, and that may dominate over the level of education.  
Not completing a year at any level may have more to do with socio-economic forces, with more poor people starting but not completing a level of education. At higher levels dropping out becomes more common. The ABD (all but dissertation) PhD candidate is all too common. Sometimes life gets in the way.

---

<div class="post-metadata">

**Author:** ![Tim\_T-Bonham.net](https://avatars.discourse-cdn.com/v4/letter/t/46a35a/32.png) [@Tim\_T-Bonham.net](https://boards.straightdope.com/u/Tim_T-Bonham.net)\
**Post date:** [September 8, 2026, 4:46am UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/5 "2026-09-08T04:46:42Z")

</div>

> [@hideousidiot](#):
>
> My question is: assuming we could gather data at each of these levels of education, would they eventually show a smooth correlation between the sixty education-levels as there is between the six education-levels we now measure?

My guess would be “No”.  
Because at that detailed a level, so much data would contain so much differences that the ‘noise’ would overwhelm the data. For example, a couple of personal items:

- At the one-room rural school I first attended, clearly un-ready students were advanced to the next grade automatically, when they should have been held back. The teacher knew this, but was forbidden from doing so by the school board (3 of the 5 were related, and greatly object to their kids, relatives kids, or neighbors kids being held back. But when they reached an age where they were fully useful on the farm, they were allowed to drop out. So that school would show a lot who left at 6th or 7th grade level, but educationally were really 4-5 years below that.)
- And in the town high school, good athletes were kept in school until graduation, because our teams needed them. So they would show up as High School graduates in this research, but educationally were barely Jr High level.
- Also, it was well known in town that the parochial school teachers were much tougher graders than at the public school, and tighter discipline (probably because they could expel or push out troublesome students).

I would expect such ‘special’ cases are common all over this country, and other countries, so that any completely ‘smooth correlation’ is unlikely.  
But I do agree that the general trend of more education = more liberal social views is roughly true.

---

<div class="post-metadata">

**Author:** ![thelurkinghorror](https://avatars.discourse-cdn.com/v4/letter/t/7c8e57/32.png) [@thelurkinghorror](https://boards.straightdope.com/u/thelurkinghorror)\
**Post date:** [September 8, 2026, 3:42pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/6 "2026-09-08T15:42:54Z")

</div>

Now you have 60 discrete, categorical, levels of education, but is it safe to say these are equally spaced or weighted? Probably not, then correlation isn’t a good measure.

Not sure what “glitches” implies, you would need enough subjects to establish sufficient [power](https://en.wikipedia.org/wiki/Power_(statistics)) to find any true differences. You would need more subjects, but not 10x as much. Software like [G\*Power](https://www.psychologie.hhu.de/arbeitsgruppen/allgemeine-psychologie-und-arbeitspsychologie/gpower) can calculate this. Adding this many comparisons would also mean that any statistical change is likely to find at least one difference by chance that doesn’t exist, so you’d need to control for that.

---

<div class="post-metadata">

**Author:** ![Pasta](https://avatars.discourse-cdn.com/v4/letter/p/ecccb3/32.png) [@Pasta](https://boards.straightdope.com/u/Pasta)\
**Post date:** [September 8, 2026, 7:23pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/7 "2026-09-08T19:23:15Z")

</div>

> [@hideousidiot](#):
>
> assuming we could gather data at each of these levels of education, would they eventually show a smooth correlation between the sixty education-levels as there is between the six education-levels we now measure?

Political leanings and gross education levels correlate with a lot more than each other. A person’s parents’ political views, their community’s political views, how urban-vs-rural their upbringing and subsequent environment, the size and diversity of their school(s), the occupations they’ve held and those near them have held, their financial well-being and that of their parents, etc., etc., all are major correlates. This is not something for which there is a simple relationship that can be zoomed in on. If you try to pick a narrow slice, you have to care about that narrow slice’s details, too. As an example: For a student to stop school after 6th grade, something rather unusual has to be happening. Family trauma? Residing in an area without good truancy monitoring? Or in an area far from school? Or in a family where parents work multiple jobs? Perhaps it is the case high-population urban schools push many more unready students through than do rural schools, or maybe that’s more dominant in rural settings after all (like in @Tim_T-Bonham.net’s case). The point is simply that _these_ effects – whatever they are – will dominate and will also be correlated to political leaning. If, say, students that drop out at 6th grade mostly come from high-population schools in areas where families make use of social programs, those students are probably living in a democratic-leaning environment. Anyway, you won’t be measuring anything that has any reason to be smooth.

> [@hideousidiot](#):
>
> That is, are there enough people who’ve started fourth grade but did not finish it in the entire country to make for stable reliable polling results, even if you were able to poll each person who qualifies for incusion?

Focusing narrowly on this statistical question (and not on the systematic uncertainties related to (lack of) controlling for much more dominant effects that make the statistical question moot): If you treat this as a two-party preference, then it’s a case of “binomial statistics”. If the preference rate for party _P_ is _p_ and for party _Q_ is _q_ = 1 - _p_, and if you have _N_ people in some narrow education bin, then your uncertainty on the observed _p_ is roughly sqrt(_pq/N_). So for p = 60% and N = 10000 samples, you’d have an estimate of roughly p=60\% \pm 0.5\%.

---

<div class="post-metadata">

**Author:** ![thelurkinghorror](https://avatars.discourse-cdn.com/v4/letter/t/7c8e57/32.png) [@thelurkinghorror](https://boards.straightdope.com/u/thelurkinghorror)\
**Post date:** [September 8, 2026, 7:50pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/8 "2026-09-08T19:50:08Z")

</div>

> [@Pasta](#):
>
> Focusing narrowly on this statistical question (and not on the systematic uncertainties related to (lack of) controlling for much more dominant effects that make the statistical question moot): If you treat this as a two-party preference, then it’s a case of “binomial statistics”. If the preference rate for party _P_ is _p_ and for party _Q_ is _q_ = 1 - _p_, and if you have _N_ people in some narrow education bin, then your uncertainty on the observed _p_ is roughly sqrt(_pq/N_). So for p = 60% and N = 10000 samples, you’d have an estimate of roughly p=60\% \pm 0.5\% 𝑝 =60% ±0.5%.

Probably best with a Leikert style rating, e.g. mostly agree with party A / somewhat agree with A / politically neutral / somewhat agree with party B / mostly agree with party B. That’s usually the way I see it in both psych studies and Pew and other polling. Even better is if you ask questions that are from part platforms, or known to correlate with them. Or if you’re not interested in partisanship but general political scales like liberal/conservative, adjust accordingly.

Your first paragraph illustrates the difficulties of drawing valid conclusions from these data. It’s hard if you’re doing a legitimate study.

---

<div class="post-metadata">

**Author:** ![echoreply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/echoreply/32/3641_2.png) [@echoreply](https://boards.straightdope.com/u/echoreply)\
**Post date:** [September 9, 2026, 1:23am UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/9 "2026-09-09T01:23:12Z")

</div>

A robust relationship can be observed between two variables, achieved education level and political affiliation in our case, and that relationship can then be examined. Any of the reasons @Pasta gave are good explanations, and educational attainment itself can still be related to political affiliation even when accounting for all of the other variables. It is almost certainly a mix of many of those things.

In the case of educational attainment my biggest concern is that even though the slices have equal lengths, I do not expect them to have an equal relationship to the outcome variable. There may be no meaningful difference between people whose maximum grade level was 7 and whose was 8. However there may be a big difference between maximum completed grade level of 11 and 12. Similar with completing a 4 year degree and dropping out after 3 years\[1\]. Those are all the same step sizes, but the steps have importance beyond their size.

In my experience that is one of the reasons that educational attainment is going to be coded as no high school, some high school, completed high school, some college, etc.

Maximum completed grade level is a simple enough example that we can all understand why 7 vs. 8 is different than 11 vs 12. For other measures it might not be as clear, and is one of the reasons that it is best to consult with subject area experts when using a measure you don’t know much about.

One silly example from my research many years ago was looking at number of joints smoked per day\[2\]. Pot heads I talked to explained that at some point “you’re just burning your stash, man” and that someone who reports smoking 8 joints per day is not getting any higher than someone who reports 4. The appropriate way to analyze the data is to add a ceiling, so anything higher than 5 is treated as 5\[3\].

* * *

1. I know it can take 6 years to complete a 4 year degree, or 3 years 

2. so long ago people still smoked joints 

3. or whatever the numbers were

---

<div class="post-metadata">

**Author:** ![hideousidiot](https://avatars.discourse-cdn.com/v4/letter/h/858c86/32.png) [@hideousidiot](https://boards.straightdope.com/u/hideousidiot)\
**Post date:** [September 10, 2026, 12:55pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/10 "2026-09-10T12:55:11Z")

</div>

Let’s try this from another perspective. Suppose you were to divide voters into only two categories: “did not graduate from high school” and “graduated from high school,” and further suppose that the first group significantly tended towards voting Republican and the second tended to vote Democratic. You’d have large sample sizes to work from, and clear results from your poll.

I don’t think Tim\_T-Bonham.net’s objections would be made. The sample sizes would be so huge that Tim\_T-Bonham.net ‘s objections would get swallowed up as rounding errors, and we could safely dismiss them.

Next, we’d (re-)introduce our original six categories, and I still don’t think those objections would be carrying much weight.

So the question now is posed as: could we expand the number of gradations of education levels to eight, or to ten, and still retain the same degree of sample-size stability and the same (or similar) outcome of steadily increasing tendency to vote Democratic as the education level rises?

I would think the answer would be: “Sure, we could.”

So my ultimate question now is: “How much could we increase that number of gradations before this sort of poll becomes unreliable?”

Again, I’ll remind you that no one seems to think that six gradations yields unreliable results, and six is a pretty arbitrary number.

As to Francis\_Vaughn’s point, I’m not at all sure that the reasons people drop out of the education system has much bearing on my question. Someone who stops attending school after the eighth grade because of economic pressure still has the same level of education as someone who drops because the work is too demanding. Is it that important whether one person drops out because a parent demands that the kid bring in needed income or because the kid joins a street gang? If that were so, we’d have the same issue with the current six-levels of education poll I referred to in the OP. The only way I can see “reasons kids drop out” being relevant is if we were to include “reasons” in the poll and thus create smaller and smaller sample sizes, but I think if we stick to clear temporal gradations, and ignore the reasons (which are subjective by their nature) we’d get clearer results.

---

<div class="post-metadata">

**Author:** ![Jasmine](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jasmine/32/2964_2.png) [@Jasmine](https://boards.straightdope.com/u/Jasmine)\
**Post date:** [September 10, 2026, 1:25pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/11 "2026-09-10T13:25:23Z")

</div>

> [@hideousidiot](#):
>
> According to an AI I consulted, and my own sense, education-levels in this country correlate closely with voting patterns: people with doctorates tend to vote Democratic much more than people who never graduated high school,

In statistics, a meaningful relationship between variables is determined by various different metrics which are dependent upon the test being used. Another factor, which you mentioned, is the size of the sample being examined. One tool is the Pearson’s r, which is a statistical measure that shows the strength and direction of a linear relationship between two quantitative variables. The point being that this is measurable.

---

<div class="post-metadata">

**Author:** ![Pasta](https://avatars.discourse-cdn.com/v4/letter/p/ecccb3/32.png) [@Pasta](https://boards.straightdope.com/u/Pasta)\
**Post date:** [September 10, 2026, 4:18pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/12 "2026-09-10T16:18:12Z")

</div>

> [@hideousidiot](#):
>
> The only way I can see “reasons kids drop out” being relevant is if we were to include “reasons” in the poll and thus create smaller and smaller sample sizes

That’s exactly what you have to do if the intention is to understand if (and how much) education level influences political leaning. The key word is “confounding variables”, and in this story there are boat loads of these that would need to be controlled for.

> [@hideousidiot](#):
>
> If that were so, we’d have the same issue with the current six-levels of education poll I referred to in the OP.

We _do_ have the same issue with the six levels, or even two levels. Social science is hard. As a greatly reduced example to emphasize the point:

Say I measure if education level relates to whether people own a truck. Result: _Oh look! Those with a college degree own way fewer trucks!_

If someone has a college degree, they more likely live in cities, more likely work office jobs, etc. Vehicle choice isn’t _caused_ by the number of classes they took or how educated they are; it’s an effect of many things that _also_ relate to education level. Saying something like “People with less education choose trucks more often” is exceedingly misleading, and depending on intent, just wrong because I haven’t controlled for confounding variables like geographic location or occupation – things that also relate to both education level and truck ownership.

So, I do a better study. I recognize that at least these two things are confounding, so I measure truck ownership strictly within Dallas County and only for those who work in the insurance industry. I split _that_ group into “college degree” or “no college degree”. Result: _Oh look! There is no significant difference at all in rate of truck ownership with education level!_

Other potential examples (I’m making them up, but they are valid pedagogically):

- People with college degrees are more likely to own the latest smartphone. However, people that never went to college will have a higher average age (college attendance has increased over the decades), and older people don’t see a need to upgrade their phones routinely. So maybe phone upgrading isn’t related to education level at all, or maybe in the opposite direction. We can’t say, because age is a big factor and age is echoed in the education level groupings, and we haven’t controlled for age.

- People with college degrees are more likely to eat at ice cream shops. But, people who have jobs in walkable or compact cities are more often in careers that require college, and their kids’ schools are more likely in cities, and these are exactly the cohorts that will be passing by ice cream shops. So, “eating at ice cream shops” is affected directly by the urban environment. We know nothing about whether education level further affects the choice to visit ice cream shops, because any such “signal” in the data is overwhelmed by the education-related confounding factor of “environment”.

> [@hideousidiot](#):
>
> “How much could we increase that number of gradations before this sort of poll becomes unreliable?

“Unreliable” is the key word. What are you trying to measure? If it’s “does more education make people more likely to vote one way or another”, then you must control for confounding variables to make any sort of claim that sounds like that. More or fewer gradations doesn’t alter this fundamental limitation. Many things make people vote how they do. And many things make people end up with different levels of education.

This is all why I and @Tim_T-Bonham.net and others were pointing out other confounding variables. For someone to drop out of school early, there are huge and non-ignorable other effects present for them that are surely drivers for their world view, much more than the fact that they never took trigonometry. And because the dominant confounding factors will be very different for the different education level gradations, there is no reason to expect the aggregate result to be always increasing with education level as you divide things up further.

(For instance, take the ice cream example again. Let’s say that there actually is an education-related _decrease_ in ice cream consumption. More education → less ice cream eating, maybe because of some “eat healthy” college classes or something. In any case, assume that it’s truly there. Now, if we do an _uncontrolled_ study, we would see the wrong, opposite effect overwhelmingly (more education → more ice cream eating), because the relationship between city environment and education level is also present and strong.)

---

<div class="post-metadata">

**Author:** ![echoreply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/echoreply/32/3641_2.png) [@echoreply](https://boards.straightdope.com/u/echoreply)\
**Post date:** [September 10, 2026, 7:12pm UTC](https://boards.straightdope.com/t/statistics-and-sample-size-question/1032882/13 "2026-09-10T19:12:15Z")

</div>

> [@hideousidiot](#):
>
> So my ultimate question now is: “How much could we increase that number of gradations before this sort of poll becomes unreliable?”

This can get into the difference between continuous variables and discrete (or categorical) variables, and the problems with both turning a continuous variable into a discrete one, and treating discrete variables as continuous. These are things a statistician needs to pay attention to so the proper analysis are performed and interpreted correctly.

Generally there are problems when categorical values are treated as continuous. It can impose false linear relationships. For example, what if people who only finished high school are more likely to vote Republican, college grads are more likely to vote Democratic, but PhD holders are more likely to vote Republican? Because there are fewer PhD holders than the other categories, and education is treated as a continuous variable, we might interpret it that more education means more likely to vote Democrat when that is not actually the case.

Information is lost when taking a continuous measure and reducing it to a few categories. For example, think if the amount of information provided by these two variables: `Age` (a number), and `Adult` (yes/no, for age under/over 18). Turning a continuous variable into a categories could cause a real effect to be missed, because statistical power is lost.

So for your example, knowing exactly how much school each subject has completed is providing more information to the poll results, not less. If the outcome is very sensitive to educational level, then it might be making the results more reliable, not less reliable.
