(I seem to get in an awful lot of arguments about probabilities and coin flips. If this post doesn’t make you sick of me, I spoke about this at further length in this thread, which may be informative.)
[QUOTE=Left Hand of Dorkness]
I got my number by finding an applet that allows you to put in two numbers, X and Y, and get the result of X^Y. If my number is screwy, I blame the applet :).
And I think we agree on what’s tested, although I may have said it poorly. It’s tremendously unlikely that the coin is fair, if 1,000 flips come up universally heads; given a fair coin, that’s an astronomically unlikely result. The two statements are almost certainly not going to be paired together.
[/QUOTE]
I agree with the statement “given a fair coin, [1000 all heads is] an astronomically unlikely result.” I feel that is pretty much tautologous; the definition of a coin being fair is such that we know P(1000 all heads | the coin is fair) = 1/2^1000.
However, without further assumption, I would not agree with the statement “It’s tremendously unlikely that the coin is fair, if 1,000 flips come up universally heads”. On what grounds? Supposing I flipped a coin 1000 times and got the precise result HTHHHTHHTHTHT…TTTHTH. This result also has probability 1/2^1000 of occurring, given a fair coin; does that mean this result is also evidence that the coin I flipped was not fair? We would be led to absurdity; every possible sequence of 1000 flips of a coin would cause us to determine that it wasn’t fair. We never think to ourselves “Oh, yes, this coin, when flipped 3 times, came up HTH. I guess that means it has probability 1/8 of being fair”, because, indeed, such reasoning is not probabilistically sound.
That is to say, without some special assumptions, it’s entirely unclear to me what P(the coin is fair | 1000 all heads) is. I can expand it out a bit, to P(the coin is fair)/[P(the coin is fair) + 2^1000 * P(the coin is unfair and 1000 heads come up)], where P(the coin is fair) is the unconditional a priori probability of a fair coin and similarly for the other such term, but without any assumptions about the a priori probabilities of the coin being fair, of the coin giving 1000 heads if it’s not fair, etc., it’s impossible to evaluate this term and conclude that it’s small, large, whatever.
Now, as it happens, I do, as a human being, make further assumptions. To the extent that it’s meaningful to speak of some coins as fair and some coins as unfair*, I happen to assume an unfair coin is more likely to give 1000 heads than to give, say, heads on every prime numbered flip below 1000 and tails on the others. And I also happen to assume that a coin is reasonably likely to be unfair, so that the second term in the denominator above is not too small. Quantifying and putting these assumptions together, I could indeed determine that P(the coin is fair | the first 1000 flips are all heads) is tiny. But those assumptions definitely need to be made and noted.
*: Presumably, the fairness or unfairness of a coin is not to be found in its flip sequence alone, a fair coin being capable of giving whatever sequence you name, and similarly for an unfair coin. So, then, what is it about a coin that constitutes its fairness? [It’s a similar situation with predicates like “X has a probability exactly 0.7 of having brown eyes.” Does this predicate hold of anyone? How could we know? All we know of a person is that they either really do have brown eyes or really don’t]. The fairness of a coin could just be a hidden unobservable variable in the universe, that some coins have set to FAIR and others to UNFAIR, but that would be an odd entity to introduce. What we really mean when we speak of some coins as unfair is something causal; we are saying that the coin has a weight upon it, or is controlled by magnets, or this or that; that it has some concrete properties which, according to the laws of physics, make it more likely to assume some flip patterns than others. But as the laws of physics are themselves inductively derived, this begins to complicate the reasoning above, in terms of avoiding circularity.
At any rate, it would simplify things if, rather than discussing P(the coin is fair | the first 1000 flips were heads), we discussed P(the next flip will be heads | the first 1000 flips were heads), which avoids the whole problem of the legitimacy of second-order probability. Nothing from above is essentially changed; in order to conclude that P(the next flip will be heads | the first 1000 flips were heads) is high, we must have had an a priori assumed probability distribution on flip sequences where P(1001 heads) was much higher than P(1000 heads followed by a tail). Making such an assumption is essentially the same thing as assuming the legitimacy of inductive argument, and thus not unreasonable, but we should be aware of what exactly we are doing, so that we don’t conflate it with things it’s not [we’re not using purely mathematical probabilistic argument alone, and we certainly need to avoid the mistake of simply equating the values of P(A|B) and P(B|A)], and, furthermore, so that we are aware of what exactly are the potential pitfalls of what we are doing and how best to grapple with them [e.g., why should we assume P(1001 heads) >> P(1000 heads followed by a tail)? Isn’t the latter just as much some kind of pattern as the former? What makes some patterns better for inductive extrapolation than others, such that we would readily rate the pattern “Heads on every even-numbered flip” as much more likely than the pattern “Heads on every even-numbered flip below 708 and every odd-numbered flip above it”?]