ComponentCountTest

class ComponentCountTest(val hypothesisedNumComponents: Int, val alternativeNumComponents: Int, val numObservations: Int, val improvementStatistic: Double?, nullStatistics: DoubleArray, val numReplicates: Int, val level: Double = defaultLevel)(source)

A procedure-calibrated test of a hypothesised component count against one more than itself.

This is not a likelihood-ratio test, and the distinction is not pedantic. This library does not maximise the likelihood: it partitions the sorted data, fits a family per group, and refines the cut positions. The log-likelihood it reports is therefore the value achieved by that procedure, not the supremum over the k-component family, so the classical statistic is not what is being computed and no textbook null applies to it. Mixture likelihood ratios are in any case non-regular — under a k-component null the larger model reaches the truth on a set rather than a point, the Fisher information is singular, and there is no chi-square limit (Hartigan 1985; Dacunha-Castelle and Gassiat 1999; Liu and Shao 2003).

The bootstrap is what makes the comparison legitimate. The null is generated by running the same procedure on data drawn from the fitted hypothesised model, so whatever this procedure falls short of an ideal maximiser by is present on both sides of the comparison and cancels. What the result means is therefore precise and narrow: is the improvement this procedure gets from an extra component larger than the improvement it gets from an extra component on data that genuinely has the hypothesised number? That is the analyst's question. It is not a claim about any ideal estimator, and it should not be reported as one.

The statistic is twice the gain in log-likelihood from the extra component, which is the quantity ComponentCountEvidence already reports per count as its marginal gain. The difference is what it is compared against: that class charges a fixed algebraic penalty which knows nothing about the sample, while this one builds the gain's own reference distribution from the data.

Parameters

hypothesisedNumComponents

the count under test

alternativeNumComponents

the count it is tested against, one greater

numObservations

the size of the sample tested

improvementStatistic

twice the observed gain in log-likelihood, or null when either fit was unusable

nullStatistics

the same quantity on each usable replicate drawn from the fitted hypothesised model

numReplicates

how many replicates were attempted, usable or not

level

the level at which the verdict is read

Constructors

Link copied to clipboard
constructor(hypothesisedNumComponents: Int, alternativeNumComponents: Int, numObservations: Int, improvementStatistic: Double?, nullStatistics: DoubleArray, numReplicates: Int, level: Double = defaultLevel)

Types

Link copied to clipboard
object Companion

Properties

Link copied to clipboard
Link copied to clipboard

Whether the bootstrap has the resolution to reject at this level at all.

Link copied to clipboard
Link copied to clipboard
Link copied to clipboard

The replicate statistics, in ascending order. A defensive copy, so that a caller cannot reorder the null out from under the quantiles.

Link copied to clipboard
Link copied to clipboard
Link copied to clipboard

How many replicates were attempted and produced nothing, whether refused before fitting or fitted to an unusable log-likelihood.

Link copied to clipboard

How many replicates produced a usable statistic.

Link copied to clipboard

The bootstrap p-value, or null when there is no statistic or no null to read it against.

Link copied to clipboard

The smallest p-value this many replicates could produce, whatever the data say.

Link copied to clipboard

The share of attempted replicates that produced a statistic.

Link copied to clipboard

The reading, at the stated level.

Functions

Link copied to clipboard

What must accompany any display of this verdict.

Link copied to clipboard

A sentence an analyst can read.

Link copied to clipboard
fun nullQuantile(proportion: Double): Double?

The replicate statistic at the given proportion, by the nearest-rank definition, or null when no replicate was usable.

Link copied to clipboard

The quantiles worth showing beside a verdict: the middle of the null and its upper tail.

Link copied to clipboard
open override fun toString(): String