Component Count Test
A procedure-calibrated test of a hypothesised component count against one more than itself.
This is not a likelihood-ratio test, and the distinction is not pedantic. This library does not maximise the likelihood: it partitions the sorted data, fits a family per group, and refines the cut positions. The log-likelihood it reports is therefore the value achieved by that procedure, not the supremum over the k-component family, so the classical statistic is not what is being computed and no textbook null applies to it. Mixture likelihood ratios are in any case non-regular — under a k-component null the larger model reaches the truth on a set rather than a point, the Fisher information is singular, and there is no chi-square limit (Hartigan 1985; Dacunha-Castelle and Gassiat 1999; Liu and Shao 2003).
The bootstrap is what makes the comparison legitimate. The null is generated by running the same procedure on data drawn from the fitted hypothesised model, so whatever this procedure falls short of an ideal maximiser by is present on both sides of the comparison and cancels. What the result means is therefore precise and narrow: is the improvement this procedure gets from an extra component larger than the improvement it gets from an extra component on data that genuinely has the hypothesised number? That is the analyst's question. It is not a claim about any ideal estimator, and it should not be reported as one.
The statistic is twice the gain in log-likelihood from the extra component, which is the quantity ComponentCountEvidence already reports per count as its marginal gain. The difference is what it is compared against: that class charges a fixed algebraic penalty which knows nothing about the sample, while this one builds the gain's own reference distribution from the data.
Parameters
the count under test
the count it is tested against, one greater
the size of the sample tested
twice the observed gain in log-likelihood, or null when either fit was unusable
the same quantity on each usable replicate drawn from the fitted hypothesised model
how many replicates were attempted, usable or not
the level at which the verdict is read
Properties
The replicate statistics, in ascending order. A defensive copy, so that a caller cannot reorder the null out from under the quantiles.
How many replicates were attempted and produced nothing, whether refused before fitting or fitted to an unusable log-likelihood.
The smallest p-value this many replicates could produce, whatever the data say.
The share of attempted replicates that produced a statistic.
The reading, at the stated level.