# Questions about making a ranking from submitted lists

**URL:** https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021
**Category:** Factual Questions
**Created:** [January 6, 2020, 8:41am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021 "2020-01-06T08:41:33Z")
**Posts on this page:** 12
**Page:** 1

<div class="post-metadata">

### Author: ![Art\_Rock](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/art_rock/32/525_2.png) [@Art\_Rock](https://boards.straightdope.com/u/Art_Rock)
#### Post date: [January 6, 2020, 8:41am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/1 "2020-01-06T08:41:33Z")

</div>

For a classical music board I’ve prepared a ranking of composers. In total 57 members sent in their top 30 composers, which I rated (40 points, gradually decreasing to 6 points, initially a bit steeper) and combined to a list. I chose a cut-off point: to be ranked in the final list, a composer had to be named at least 3 times. In the end, I could create a top100 this way. I realize that most exact rankings in this top100 are statistically not valid, but I have two questions.

[1] Someone objected that by using top30’s I could not go beyond a top 30 for the results. This sounds wrong to me, but I can’t find anything to disprove (or prove) the statement.

[2] Should I have gone for a different number than 3 for the cutoff for statistical reasons?

Thanks for any help.

---

<div class="post-metadata">

### Author: ![Ludovic](https://avatars.discourse-cdn.com/v4/letter/l/7ab992/32.png) [@Ludovic](https://boards.straightdope.com/u/Ludovic)
#### Post date: [January 6, 2020, 11:01am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/2 "2020-01-06T11:01:57Z")

</div>

I’m not sure about #2, but for #1, consider a very segmented polling population. Let’s round to 56 respondents and divide them into two groups of 28, and every respondent in a given group answers in the exact same way, but no one repeats a composer _from the other group_. Thus you have 60 different composers, and to me it would seem more arbitrary to cut off the results at #15 from each group rather than have them include all 60, since even the lowest composer got 168 points.

---

<div class="post-metadata">

### Author: ![septimus](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/septimus/32/410_2.png) [@septimus](https://boards.straightdope.com/u/septimus)
#### Post date: [January 6, 2020, 11:23am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/3 "2020-01-06T11:23:46Z")

</div>

We’ve had similar questions before; I’ll just make one comment here.

> [@Art\_Rock](#):
>
> [1] Someone objected that by using top30’s I could not go beyond a top 30 for the results. This sounds wrong to me, but I can’t find anything to disprove (or prove) the statement.

I like to think about “corner cases.” Suppose hypothetically that _everyone_ ranked Leonard Cohen as the #31 composer. That would mean that he “should” be ranked well above #31. But in your scheme he doesn’t even make the Top 100.

---

<div class="post-metadata">

### Author: ![Banksiaman](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/banksiaman/32/3334_2.png) [@Banksiaman](https://boards.straightdope.com/u/Banksiaman)
#### Post date: [January 6, 2020, 11:41am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/4 "2020-01-06T11:41:12Z")

</div>

Once scored, totalled and those with less than 3 votes omitted then you just do a sort on the scores?

Do you have a way to distinguish between similar scores with different characteristics? From what you describe, Dave Mozart can score say 240 points from six fan-boys who’ve scored him as top, while no-one else even put him on their lists, versus 240 from 40 different voters who were unanimous in putting him last.

If this was a proper music contest, like Eurovision, then the top 10 get scores, while lower ranked get nul points. This seems to separate the favourites from consistent also-rans much more quickly.

---

<div class="post-metadata">

### Author: ![Art\_Rock](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/art_rock/32/525_2.png) [@Art\_Rock](https://boards.straightdope.com/u/Art_Rock)
#### Post date: [January 6, 2020, 2:22pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/5 "2020-01-06T14:22:52Z")

</div>

Thanks for all reactions so far, I’m really looking for an answer from a statistics point of view. I want to counter the people stating that of course you can’t go beyond 30 in this case, and I know them - mental exercises like sketched here will be brushed aside as irrelevant.

---

<div class="post-metadata">

### Author: ![Pasta](https://avatars.discourse-cdn.com/v4/letter/p/ecccb3/32.png) [@Pasta](https://boards.straightdope.com/u/Pasta)
#### Post date: [January 7, 2020, 8:36pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/6 "2020-01-07T20:36:58Z")

</div>

I’m on board with the objection to going much beyond 30 if that’s all any one person could submit. Certainly not to 100. I’m surprised, actually, that you even have 100 unique names left after your “named three times” criterion.

Consider ice cream flavors. Ask 30 people to provide their top 6, and then try to make a list of the top 20 flavors. It’s not going to make any sense. The “true” 18th, 19th, 20th-ranked flavors should end up being kinda-weird-but-not-completely-crazy stuff like cucumber or whatever, but cucumber isn’t going to appear in anyone’s top 6, so it has no way to show up in the list where it belongs. In this example, you probably just won’t end up with 20 unique flavors to fill the list, but if you somehow managed to, the poorly ranked flavors will represent individual outliers in people’s top 6 rather than any sort of consensus opinion about what should be down at those rankings.

(To be more specific: Say cucumber should be the true 20th, but nobody puts it in their top 6 [who would?]. But, one person is really keen on pineapple and ranks it 5th, and another is really keen on blueberry and also ranks it 5th, and some weirdo puts carrot in their 6th slot. Those become the flavors that can end up in 20th place, but everyone would agree that pineapple and blueberry should be around, maybe, 12th and that carrot shouldn’t make the top 20 at all. And poor cucumber never even gets a chance!)

---

<div class="post-metadata">

### Author: ![Hari\_Seldon](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/hari_seldon/32/5173_2.png) [@Hari\_Seldon](https://boards.straightdope.com/u/Hari_Seldon)
#### Post date: [January 7, 2020, 10:02pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/7 "2020-01-07T22:02:46Z")

</div>

First off, there is actually a theorem (Arrow’s theorem) that states that no voting method can satisfy a bunch of quite reasonable looking criteria. What I would have done would have been to give each of the voters 100 points to distribute as they see fit. If they want to give Beethoven 100 points and not give anyone else any, so be it. Then just add all the point totals and take the top 100.

---

<div class="post-metadata">

### Author: ![md2000](https://avatars.discourse-cdn.com/v4/letter/m/73ab20/32.png) [@md2000](https://boards.straightdope.com/u/md2000)
#### Post date: [January 7, 2020, 10:53pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/8 "2020-01-07T22:53:49Z")

</div>

The other issue, as mentioned by Ludovic in the second post - if it’s a diverse group with two or more distinct preference groups, then a single ranking list is meaningless unless you can guarantee the population was fairly chosen. (Think “who should be president?” poll)

---

<div class="post-metadata">

### Author: ![Art\_Rock](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/art_rock/32/525_2.png) [@Art\_Rock](https://boards.straightdope.com/u/Art_Rock)
#### Post date: [January 8, 2020, 11:26am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/9 "2020-01-08T11:26:45Z")

</div>

Thanks for the reactions. Food for thought here.

---

<div class="post-metadata">

### Author: ![septimus](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/septimus/32/410_2.png) [@septimus](https://boards.straightdope.com/u/septimus)
#### Post date: [January 8, 2020, 11:58am UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/10 "2020-01-08T11:58:19Z")

</div>

> [@Pasta](#):
>
> Consider ice cream flavors. Ask 30 people to provide their top 6, and then try to make a list of the top 20 flavors. It’s not going to make any sense. The “true” 18th, 19th, 20th-ranked flavors should end up being kinda-weird-but-not-completely-crazy stuff like cucumber or whatever, but cucumber isn’t going to appear in anyone’s top 6, so it has no way to show up in the list where it belongs. In this example, you probably just won’t end up with 20 unique flavors to fill the list, but if you somehow managed to, the poorly ranked flavors will represent individual outliers in people’s top 6 rather than any sort of consensus opinion about what should be down at those rankings.

Your point is clear (but … _cucumber_?? 😛 ); let me give another real-world example:

I’ll guess relatively few people would put _Shawshank Redemption_ on their Top Five Movie list, let alone their #1 slot. Yet there it is, sitting at the very very top of IMDB’s Top 250, the only 9.2 on the list. Lots of people put _Godfather_ way ahead of _Shawshank_, but others don’t like gangster flicks. Lots put _Casablanca_ at the very top, but it’s black-white and WWII is ancient history for Millennials. Most put _Lord of the Rings_ near the top, but some don’t like fantasy. \*\* But everybody appreciates _Shawshank Redemption_ to some extent.\*\*

---

<div class="post-metadata">

### Author: ![beowulff](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/beowulff/32/542_2.png) [@beowulff](https://boards.straightdope.com/u/beowulff)
#### Post date: [January 8, 2020, 3:25pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/11 "2020-01-08T15:25:02Z")

</div>

> [@septimus](#):
>
> Your point is clear (but … _cucumber_?? 😛 ); let me give another real-world example:
> 
> I’ll guess relatively few people would put _Shawshank Redemption_ on their Top Five Movie list, let alone their #1 slot. Yet there it is, sitting at the very very top of IMDB’s Top 250, the only 9.2 on the list. Lots of people put _Godfather_ way ahead of _Shawshank_, but others don’t like gangster flicks. Lots put _Casablanca_ at the very top, but it’s black-white and WWII is ancient history for Millennials. Most put _Lord of the Rings_ near the top, but some don’t like fantasy. \*\* But everybody appreciates _Shawshank Redemption_ to some extent.\*\*

Which is why Taco Bell always wins “Best Mexican food” in readers polls out here.

---

<div class="post-metadata">

### Author: ![Pleonast](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pleonast/32/1183_2.png) [@Pleonast](https://boards.straightdope.com/u/Pleonast)
#### Post date: [January 8, 2020, 6:57pm UTC](https://boards.straightdope.com/t/questions-about-making-a-ranking-from-submitted-lists/846021/12 "2020-01-08T18:57:27Z")

</div>

> [@Art\_Rock](#):
>
> For a classical music board I’ve prepared a ranking of composers. In total 57 members sent in their top 30 composers, which I rated (40 points, gradually decreasing to 6 points, initially a bit steeper) and combined to a list. I chose a cut-off point: to be ranked in the final list, a composer had to be named at least 3 times. In the end, I could create a top100 this way. I realize that most exact rankings in this top100 are statistically not valid, but I have two questions.
> 
> [1] Someone objected that by using top30’s I could not go beyond a top 30 for the results. This sounds wrong to me, but I can’t find anything to disprove (or prove) the statement.
> 
> [2] Should I have gone for a different number than 3 for the cutoff for statistical reasons?
> 
> Thanks for any help.

Here’s how I would approach it.

1. Make a list of every composer ranked by any member.
2. Consider every possible pair of composers. What percentage of members preferred composer A over composer B? If a member ranked one composer in the pair but not the other, then the ranked composer is preferred. If a member ranked neither composer in the pair, then that member doesn’t contribute to that percentage.
3. Evaluate each composer’s percentages. Their “natural” ranking is one plus the number of their pairwise percentages less than 50%. There is likely to be ties at multiple ranks, but up to this point it’s difficult to argue the process is unfair.
4. There are many ways to break the ties. I’d use the median pairwise percentage score. The mean gives weight to the extremes, while the median estimates the center better.
5. Showing the intermediate results would probably be interesting to the members and give some transparency to the process.

Any chance you can make available your raw data, anonymized if necessary?
