MixtureModeler

class MixtureModeler(data: DoubleArray, fitter: ComponentFitterIfc = PDFComponentFitter(), val criterion: MixtureCriterionIfc = MixtureBICCriterion(), minimumGroupSize: Int = AdmissibilityCertificate.defaultMinimumGroupSize, minimumDistinctValues: Int = AdmissibilityCertificate.defaultMinimumDistinctValues)(source)

Fits a mixture of continuous distributions to univariate data by partitioning the sorted sample, fitting a component to each part, and ranking the assembled mixtures.

By default the facade reproduces the configuration the designed experiment measured as best: a Jenks initial partition, refinement by local search over the cut positions, and separable family selection. Every one of those is a parameter of fit, so the unrefined baseline and the exhaustive family search remain available for comparison — but a user who asks for nothing in particular gets the configuration the evidence supports rather than the simplest one.

The number of components to consider is intersected with what the data can structurally support before anything is fitted, so an impossible request is reported rather than discovered as a failure part-way through.

Which entry point to use. fitSpecified is the documented path and fit over a range is the sensitivity tool, which is the opposite of what a distribution-fitting facade usually implies. The reason is measured rather than stylistic: across the designed experiment the criterion recovers the true number of components on about 36% of attempts, under-selects eight times more often than it over-selects, and is confidently wrong when it is wrong — the median criterion gap to the truth is around 19, so the size of the gap is not a signal that the choice is reliable. An analyst who has a mechanism in mind, or who has read a count off a plot, holds better information than the search does, and this class should not silently overrule them.

So: name the count when you have one, and use the range form to ask how sensitive the answer is to that choice. Neither is more supported than the other computationally; the difference is which question is being asked.

Parameters

data

the observations, which need not be sorted; a sorted copy is taken

fitter

the component fitter, wrapped in a cache by this class

criterion

the criterion used to rank assembled mixtures

minimumGroupSize

the least number of observations permitted in a group

minimumDistinctValues

the least number of distinct values permitted in a group

Constructors

Link copied to clipboard
constructor(data: DoubleArray, fitter: ComponentFitterIfc = PDFComponentFitter(), criterion: MixtureCriterionIfc = MixtureBICCriterion(), minimumGroupSize: Int = AdmissibilityCertificate.defaultMinimumGroupSize, minimumDistinctValues: Int = AdmissibilityCertificate.defaultMinimumDistinctValues)

Types

Link copied to clipboard
object Companion

Properties

Link copied to clipboard

The structural facts of the sample, computed once.

Link copied to clipboard
Link copied to clipboard

A histogram of the sample.

Link copied to clipboard

The largest number of components this sample can structurally support.

Link copied to clipboard

The observations in sorted order.

Link copied to clipboard

The sample statistics.

Functions

Link copied to clipboard
fun assessFitAdequacy(baseline: MixtureModelingResults, numReplicates: Int = FitAdequacy.defaultNumReplicates, level: Double = FitAdequacy.defaultLevel, fitter: () -> ComponentFitterIfc = { PDFComponentFitter() }, refiner: () -> PartitionRefinerIfc = defaultRefiner, selector: () -> FamilySelectorIfc = defaultSelector, streamNum: Int = 0, streamProvider: RNStreamProviderIfc = KSLRandom.DefaultRNStreamProvider): FitAdequacy?

Whether the fitted density reproduces the sample it came from.

Link copied to clipboard
fun assessModality(numBootstrapSamples: Int = ModalityAnalyzer.defaultNumBootstrapSamples, streamNumber: Int = 0, streamProvider: RNStreamProviderIfc = KSLRandom.DefaultRNStreamProvider): ModalityAssessment

How many modes the data appears to have, and whether that appearance survives smoothing.

Link copied to clipboard
fun describe(numBootstrapSamples: Int = ModalityAnalyzer.defaultNumBootstrapSamples, streamNumber: Int = 0, streamProvider: RNStreamProviderIfc = KSLRandom.DefaultRNStreamProvider): MixtureDataDescription

Everything worth knowing about the sample before fitting anything to it.

Link copied to clipboard
fun fit(numComponents: Int, generator: PartitionGeneratorIfc = JenksPartitionGenerator(mySortedData, maxRequested(listOf(numComponents))), refiner: PartitionRefinerIfc = defaultRefiner(), selector: FamilySelectorIfc = defaultSelector()): MixtureModelingResults

Fits a mixture with the number of components the caller has chosen.

fun fit(numComponentsRange: Iterable<Int> = defaultNumComponentsRange, generator: PartitionGeneratorIfc = JenksPartitionGenerator(mySortedData, maxRequested(numComponentsRange)), refiner: PartitionRefinerIfc = defaultRefiner(), selector: FamilySelectorIfc = defaultSelector()): MixtureModelingResults
Link copied to clipboard
fun fitSpecified(numComponents: Set<Int>, generator: PartitionGeneratorIfc = JenksPartitionGenerator(mySortedData, maxRequested(numComponents)), refiner: PartitionRefinerIfc = defaultRefiner(), selector: FamilySelectorIfc = defaultSelector()): MixtureModelingResults

Fits the numbers of components the caller is choosing between.

Link copied to clipboard
fun partitionStability(baseline: MixtureModelingResults, displacements: List<Double> = PartitionStability.defaultDisplacements, selector: FamilySelectorIfc = defaultSelector(), gapShare: Double = PartitionStability.defaultGapShare, gapMovement: Double = PartitionStability.defaultGapMovement, followsShare: Double = PartitionStability.defaultFollowsShare): PartitionStability?

How far the reported density moves when the cuts between the components are nudged.

Link copied to clipboard
fun plausibleComponentCounts(numComponentsRange: Iterable<Int> = defaultNumComponentsRange, numReplicates: Int = ComponentCountTest.defaultNumReplicates, level: Double = ComponentCountTest.defaultLevel, fitter: () -> ComponentFitterIfc = { PDFComponentFitter() }, refiner: () -> PartitionRefinerIfc = defaultRefiner, selector: () -> FamilySelectorIfc = defaultSelector, streamNum: Int = 0, streamProvider: RNStreamProviderIfc = KSLRandom.DefaultRNStreamProvider): PlausibleComponentCounts

The component counts this sample does not rule out.

Link copied to clipboard
fun testComponentCount(hypothesisedNumComponents: Int, numReplicates: Int = ComponentCountTest.defaultNumReplicates, level: Double = ComponentCountTest.defaultLevel, fitter: () -> ComponentFitterIfc = { PDFComponentFitter() }, refiner: () -> PartitionRefinerIfc = defaultRefiner, selector: () -> FamilySelectorIfc = defaultSelector, streamNum: Int = 0, streamProvider: RNStreamProviderIfc = KSLRandom.DefaultRNStreamProvider): ComponentCountTest?

Tests a component count the analyst proposed against one more than itself.

Link copied to clipboard
open override fun toString(): String