SeparabilityBound

How far a k-component mixture is from the nearest mixture with one component fewer, and the sample size that distance implies.

Why this is worth computing. Every measured statement about component recovery so far has been about what one procedure achieves. This is about what any procedure could achieve. If the true mixture is very close to some mixture with fewer components, then no amount of cleverness in the fitting recovers the count from a small sample: the two models are nearly indistinguishable in the data, and the question is information-theoretic rather than algorithmic.

The governing quantity is

    eta(k) = inf over all (k-1)-component mixtures g of Hellinger( f(k), g )

and classical testing theory gives a necessary condition on the sample size, of the order of one over eta squared, to tell the two apart with bounded error. It is a lower bound: no procedure does better, and a real procedure that must also estimate parameters does worse.

The infimum is approximated from above. The candidate set here is every mixture obtained by merging one adjacent pair of true components into a single distribution, over a catalog of families and a refinement of that family's moments. That is a subset of all (k-1)-component mixtures, so the value returned is at least the true infimum. Since the implied sample size varies as one over the square of the distance, over-stating the distance under-states the sample size: every figure this produces is a floor, not an estimate.

Only adjacent pairs are merged. For components ordered on the line and separated in the way the design separates them, a non-adjacent merge is further away, so it cannot supply the infimum.

Types

Link copied to clipboard
data class Bound(val label: String, val hellinger: Double, val kolmogorov: Double, val mergedPair: Int, val mergedFamily: String, val mergedComponent: ContinuousDistributionIfc)

What the search found for one case.

Functions

Link copied to clipboard
fun boundFor(label: String, weights: DoubleArray, components: List<ContinuousDistributionIfc>, types: List<RVParametersTypeIfc>, numRefinements: Int = 6, numPoints: Int = 8001): SeparabilityBound.Bound?

The closest (k-1)-component mixture reachable by merging one adjacent pair.

Link copied to clipboard
fun dkwSampleSize(eta: Double, alpha: Double = 0.05): Double

The Dvoretzky–Kiefer–Wolfowitz sample size: how many observations are needed for the empirical distribution function to lie within a given distance of the truth, uniformly, with the requested confidence.

Link copied to clipboard

The mean and variance of the sub-mixture formed by two weighted components.

Link copied to clipboard

Builds a distribution of the requested family with the requested mean and variance, or null where those moments are not attainable by that family.

Link copied to clipboard
fun near(a: Double, b: Double, tolerance: Double = 1.0E-9): Boolean

Indicates whether two moment values are close enough to treat as the same.