Component Count Evidence
Why the criterion chose the number of components it chose, and how much that is worth.
The central caution, which every display of this must carry. It is tempting to read the size of the criterion's margin as a measure of how likely its choice is to be right. Measured across the designed experiment, it is not: when the criterion picks the wrong component count, the median gap to the truth is about 19 units, and only about 4.6% of wrong calls fall inside a conventional weak-evidence margin of two. A large margin therefore means the criterion is not indifferent. It does not mean the criterion is right, and presenting it as though it did would give an analyst a confident reason to defer at exactly the moment deferring is wrong.
A small margin is the half that survives measurement: it genuinely does mean the criterion is nearly indifferent between two counts, which is worth knowing.
What is worth more is agreement. The criteria fail in opposite directions for reasons that are theorems — AIC over-selects, BIC under-selects — so their concurrence carries information their individual confidence does not. On the attempts where AIC and BIC agreed, the agreed count was correct on about 66% against a base rate of 37%, and the interval between their two choices covered the truth about 78% of the time. That interval is wide, averaging around three and a half of the eight candidate counts, and at the worst separations it is close to uninformative — but a wide interval is an answer too.
On a negative marginal gain. In textbook maximum-likelihood mixture fitting the log-likelihood cannot fall as components are added, because the models nest. Here they do not: the partition is regenerated for each count rather than split from the previous one, so the search can find a worse partition at the larger count. A negative gain is a real finding about the search rather than a numerical fault, and every display of it says so.
Parameters
one row per candidate count, in increasing order
the count the analyst supplied, when they supplied one
what each reported criterion would choose
Properties
Whether every reported criterion chose the same number of components.
The range of counts the reported criteria span, or null when none of them chose.
How many candidate counts the criteria span, which is how much the interval is worth.