CountAdequacy

data class CountAdequacy(val hypothesisedCount: Int, val separability: Double, val sampleSize: Int, val verdict: AdequacyVerdict)(source)

How much data the fitted component count needs, and whether this sample has it.

The question this answers is the one that comes before the count. An analyst who supplies a component count is entitled to know whether their sample could settle it, and to be told so before being shown a number that looks as though it did.

The governing quantity is the separability of the fitted mixture — how far it sits from the nearest mixture with one component fewer, measured as a Hellinger distance. Across the designed experiment, recovery of the component count depends on the sample size and that separability only through their product n * separability^2, and does so strongly: over 11,520 attempts the count was recovered on 4.4% of attempts this rates INSUFFICIENT, 39.3% of those it rates MARGINAL, and 69.7% of those it rates AMPLE.

Four things must be said wherever this is displayed, and caveats says them:

  1. The separability is measured on the fitted mixture, not on the truth, so a poor fit gives a poor estimate. It is an order of magnitude, not a measurement.

  2. The constant relating sample size to separability was measured on one designed experiment. It is a rule of thumb whose transportability is untested.

  3. AMPLE means the information is present, not that the answer is right, and whether more data would help depends on something the analyst cannot observe. Where the fitter's catalog holds the truth's families, recovery keeps climbing with n * separability^2 and reaches about 98% at the highest ratios measured. Where it does not, recovery peaks near 55% and then falls as data accumulates, because more data is more evidence for extra components to patch a family that cannot fit. The pooled figure of 77% averages those two and describes neither.

  4. The separability is an infimum approximated from above, and the required sample size varies as one over its square, so every figure here is a floor rather than an estimate.

Parameters

hypothesisedCount

the number of components the fit used

separability

the Hellinger distance to the nearest mixture with one component fewer

sampleSize

the number of observations the fit was made from

verdict

the reading of adequacyRatio against the measured curve

Constructors

Link copied to clipboard
constructor(hypothesisedCount: Int, separability: Double, sampleSize: Int, verdict: AdequacyVerdict)

Types

Link copied to clipboard
object Companion

Properties

Link copied to clipboard

The quantity recovery actually follows: the sample size times the squared separability.

Link copied to clipboard
Link copied to clipboard
Link copied to clipboard
Link copied to clipboard

The sample size at which recovery of this count reaches even odds, of the order of the measured constant over the squared separability.

Link copied to clipboard

Functions

Link copied to clipboard

The four things that must accompany any display of this verdict.

Link copied to clipboard

A sentence an analyst can read, with the verdict's own limitation attached to it.