FitAdequacy

class FitAdequacy(val numObservations: Int, val observedStatistic: Double?, nullStatistics: DoubleArray, val numReplicates: Int, val functionals: List<FunctionalAgreement> = emptyList(), val level: Double = defaultLevel)(source)

Whether a fitted mixture reproduces the sample it came from.

A different question from CountAdequacy, and usually the more important one. That class asks whether the sample can settle how many components there are. This asks whether the density is a fair description of the data — which is what an input model is for. A mixture can be thoroughly wrong about its components while reproducing the density well enough to drive a simulation, and the designed experiment says that is the common case.

Three readings are offered because they answer to different circumstances.

The bootstrap-calibrated test uses all the data and is correctly calibrated. A plain goodness-of-fit p-value against the fitted density is optimistic: the density was estimated from the very data being tested, so it sits closer to them than a distribution specified in advance would. The bootstrap removes that by refitting on every replicate, so the same over-fitting is present on both sides of the comparison and cancels.

Comparisons against held-out data were tried and dropped. They look like they should need no calibration, since a density fitted to one sample is independent of another, and that reasoning is incomplete. Holding data out removes the over-fitting — the density is not tuned to those particular observations — but not the estimation error: the fitted density is still not the truth, and the standard Kolmogorov-Smirnov null assumes a fully specified distribution rather than an estimated one. Measured where the null was true by construction, a one-sample test of held-out data against the fit rejected 13.5% of the time at a stated 5%, and a two-sample test against a draw from the fit rejected 37.5%. The bootstrap rejected 3.6%. Only the bootstrap carries the estimation error on both sides of the comparison, which is why only it is offered.

A two-sample test between the fitting data and a draw from the fit fails in the other direction, by the same reasoning rather than by measurement: the draw inherits whatever the fit took from those data, so the statistic comes out too small and the test under-rejects — conservative in the one direction an adequacy check cannot afford. That arrangement was never measured here, so treat the direction as argued and the size as unknown. Note that it is not the 37.5% reading above, which used held-out data and therefore over-rejected.

Overlaying those same two samples as a picture is a different matter and is worth doing: FittedAgainstDataPlot draws it. A figure claims no level, so the dependence that invalidates the test does not spoil the display.

Parameters

numObservations

the size of the fitting sample

observedStatistic

the Kolmogorov-Smirnov distance between the sample and its own fit

nullStatistics

the same distance on each usable replicate drawn from the fit and refitted

numReplicates

how many replicates were attempted

functionals

summaries of the data beside the same summaries of the fit

level

the level at which the verdict is read

Constructors

Link copied to clipboard
constructor(numObservations: Int, observedStatistic: Double?, nullStatistics: DoubleArray, numReplicates: Int, functionals: List<FunctionalAgreement> = emptyList(), level: Double = defaultLevel)

Types

Link copied to clipboard
object Companion

Properties

Link copied to clipboard

Whether the bootstrap has the resolution to reject at this level at all.

Link copied to clipboard
Link copied to clipboard

The largest relative error over the functionals that have one, or null when none does.

Link copied to clipboard
Link copied to clipboard

The replicate distances, ascending. A copy, so the quantiles cannot be reordered.

Link copied to clipboard
Link copied to clipboard
Link copied to clipboard

How many replicates produced a usable distance.

Link copied to clipboard
Link copied to clipboard

The bootstrap p-value, in the form that cannot return zero from a finite number of replicates, or null when there is nothing to compare.

Link copied to clipboard

The smallest p-value this many replicates could produce, whatever the data say.

Link copied to clipboard

The share of attempted replicates that produced one.

Link copied to clipboard

The reading, at the stated level. INCONCLUSIVE when the bootstrap could not reach the level however the data fell, or when too many replicates failed for the survivors to stand for the null.

Functions

Link copied to clipboard

What must accompany any display of this verdict.

Link copied to clipboard

A sentence an analyst can read.

Link copied to clipboard
fun nullQuantile(proportion: Double): Double?

The replicate distance at the given proportion, by nearest rank, or null when none.

Link copied to clipboard
open override fun toString(): String