Fit Adequacy
Whether a fitted mixture reproduces the sample it came from.
A different question from CountAdequacy, and usually the more important one. That class asks whether the sample can settle how many components there are. This asks whether the density is a fair description of the data — which is what an input model is for. A mixture can be thoroughly wrong about its components while reproducing the density well enough to drive a simulation, and the designed experiment says that is the common case.
Three readings are offered because they answer to different circumstances.
The bootstrap-calibrated test uses all the data and is correctly calibrated. A plain goodness-of-fit p-value against the fitted density is optimistic: the density was estimated from the very data being tested, so it sits closer to them than a distribution specified in advance would. The bootstrap removes that by refitting on every replicate, so the same over-fitting is present on both sides of the comparison and cancels.
Comparisons against held-out data were tried and dropped. They look like they should need no calibration, since a density fitted to one sample is independent of another, and that reasoning is incomplete. Holding data out removes the over-fitting — the density is not tuned to those particular observations — but not the estimation error: the fitted density is still not the truth, and the standard Kolmogorov-Smirnov null assumes a fully specified distribution rather than an estimated one. Measured where the null was true by construction, a one-sample test of held-out data against the fit rejected 13.5% of the time at a stated 5%, and a two-sample test against a draw from the fit rejected 37.5%. The bootstrap rejected 3.6%. Only the bootstrap carries the estimation error on both sides of the comparison, which is why only it is offered.
A two-sample test between the fitting data and a draw from the fit fails in the other direction, by the same reasoning rather than by measurement: the draw inherits whatever the fit took from those data, so the statistic comes out too small and the test under-rejects — conservative in the one direction an adequacy check cannot afford. That arrangement was never measured here, so treat the direction as argued and the size as unknown. Note that it is not the 37.5% reading above, which used held-out data and therefore over-rejected.
Overlaying those same two samples as a picture is a different matter and is worth doing: FittedAgainstDataPlot draws it. A figure claims no level, so the dependence that invalidates the test does not spoil the display.
Parameters
the size of the fitting sample
the Kolmogorov-Smirnov distance between the sample and its own fit
the same distance on each usable replicate drawn from the fit and refitted
how many replicates were attempted
summaries of the data beside the same summaries of the fit
the level at which the verdict is read
Constructors
Properties
The largest relative error over the functionals that have one, or null when none does.
The replicate distances, ascending. A copy, so the quantiles cannot be reordered.
The smallest p-value this many replicates could produce, whatever the data say.
The share of attempted replicates that produced one.
The reading, at the stated level. INCONCLUSIVE when the bootstrap could not reach the level however the data fell, or when too many replicates failed for the survivors to stand for the null.