# Checking Parity of Data Sets?

**URL:** <https://boards.straightdope.com/t/checking-parity-of-data-sets/972412>\
**Category:** Factual Questions\
**Created:** [September 28, 2022, 9:10pm UTC](https://boards.straightdope.com/t/checking-parity-of-data-sets/972412 "2022-09-28T21:10:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sage\_Rat](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/sage_rat/32/399_2.png) [@Sage\_Rat](https://boards.straightdope.com/u/Sage_Rat)\
**Post date:** [September 28, 2022, 9:10pm UTC](https://boards.straightdope.com/t/checking-parity-of-data-sets/972412/1 "2022-09-28T21:10:20Z")

</div>

I have a complete set of data Q. From Q, we calculate a bunch of new data but, previous to doing so, we drop about 4% of all information that we’ve probabilistically determined is spam.

We have developed a new set of tools that will perform these calculations, and which also drops about 5% of the data but uses a different method of spam detection. In general, we expect that there’s about 1% difference in what is in and not in the data the old way threw out slightly different things than the new way does. Much will overlap but some will not.

We want to ensure that our new technology is producing the same results as the old tools but, because the data being dropped is not quite the same, we know that there will be some variance between the two sets. More importantly, we know that values that are less common in the results will vary by a greater amount,

For example, if there were 2 penguin owners in a population of a million people, in the old result, then having that go up to 3 or down to 1 is a giant jump relative to the original. The data might have only changed by 1% but our penguinOwnerCount has changed by 50%.

Is there a formula for saying that if there was a 1% change in the dataset then, based on the size of a particular subset as a proportion of the whole, we should expect the new subset size to be within bounds M and N?

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [September 28, 2022, 11:24pm UTC](https://boards.straightdope.com/t/checking-parity-of-data-sets/972412/2 "2022-09-28T23:24:25Z")

</div>

Perhaps one of the statistical location tests listed on this page?

> **[Location test](https://en.wikipedia.org/wiki/Location_test)**
>
> A location test is a statistical hypothesis test that compares the location parameter of a statistical population to a given constant, or that compares the location parameters of two statistical populations to each other. Most commonly, the location parameter (or parameters) of interest are expected values, but location tests based on medians or other measures of location are also used.
> The one-sample location test compares the location parameter of one sample to a given constant. An example of...
