Counting One Witness Four Times
The previous post ended on an admission: the running-product construction we use to accumulate evidence is valid when the pieces are independent, and our pieces are not. This post is about what we did with that, because it is the part of the design that required real work rather than reading.
The failure, stated plainly
We run a set of detectors over each company. Several distress models — Ohlson, Altman, Zmijewski, and the hazard-model family — take overlapping inputs and are, in substance, different arithmetic on the same underlying financial position. Several accrual measures likewise: Beneish and Sloan both key on accruals, from different directions.
Now consider a company in genuine financial difficulty. Every distress model fires. If we multiply those together as independent evidence, the product explodes — and it explodes not because we have found four separate reasons for concern, but because we asked one question four times and wrote down four yeses.
The mathematics is unforgiving here. The whole guarantee rests on the expectation of the product being bounded when the null is true. Under positive dependence that bound fails. The number still goes up, the threshold still gets crossed, the flag still gets raised — and the false-positive rate it claims to respect is fiction. This is the most dangerous class of bug we deal with, because the system looks like it is working harder when it is actually just double-counting.
The two ways to combine, and what each costs
There are two elementary facts about combining these evidence measures, and the design falls out of the tension between them.
The product is the powerful one. Under independence it extracts the most information from a set of observations; that multiplicative behaviour is exactly what lets weak signals from genuinely separate sources compound into something decisive. It is also the one that breaks under correlation.
The average is the safe one. The expectation of an average is the average of the expectations, and that holds no matter how the terms relate. Arbitrary dependence, any correlation structure, does not matter. It is valid always — and it is weak, because averaging deliberately discards the compounding that made the product worth having.
Choosing one globally is choosing between a system that lies and a system that cannot see. So we do not choose globally.
Clusters inside, product across
The structure we landed on: average within a theme, multiply across themes.
We estimate the correlation structure across the detector set on the market as a whole — after transforming each feature to its position in the market distribution, so the estimate is about co-movement rather than scale. Hierarchical clustering on that structure separates the detectors into thematic groups: a distress cluster, an accrual-manipulation cluster, a market-structure cluster, and so on. Within a cluster, correlation is high by construction. Across clusters, it is low.
Then the two combination rules go exactly where each is honest. Inside a cluster we average — arbitrary dependence is fine, and four correlated distress models contribute one theme’s worth of evidence instead of four. Across clusters we multiply — these are close to independent lines of enquiry, and multiplying is what turns “unusual in distress terms and unusual in accrual terms and unusual in market-structure terms” into something that stands out sharply from a company that is merely one of those.
This is not a statistical trick bolted onto the architecture. It is the architecture written in arithmetic. What we always claimed to do was accumulate independent lines of forensic evidence. The clustering is the step that makes the claim literally true instead of aspirational — because before it, “independent lines” was a description of how we thought about the detectors, not a property of how they were combined.
Where it is still approximate, and why we say so
Three places, and none of them is fixable by being more clever:
The clusters come from an estimate. The correlation structure is measured on the market, and measured things have error. Cross-cluster independence is close, not exact. So the guarantee is an approximation, not a theorem about our actual running system.
The correlation estimate itself makes assumptions. Rank-based transformation buys robustness against heavy tails, which financial data has in abundance, but estimating a dependence structure is still estimation.
A theme can be split wrong. Put two genuinely correlated detectors in different clusters and you are back to double-counting, quietly.
Two things follow. First, we partition conservatively — when a detector is ambiguous, we would rather split a cluster too finely than merge two things that turn out to be the same question. Over-splitting costs statistical power; under-splitting costs correctness, and only one of those produces a false accusation.
Second, and more important: we do not defend the false-positive rate by argument. We measure it. A calibration test runs synthetic companies with a known dependence structure through the actual combination path and reports the observed false-positive rate against the level we claim. If the measured rate exceeds the claim, the clustering is wrong and the scoring does not ship — regardless of how good the reasoning sounded.
That check exists because the reasoning above is exactly the kind that feels airtight and can still be wrong in a way no amount of re-reading reveals. The assumption we cannot verify by thinking is the one that has to be measured.
All figures are system-level results current as of publication date; methodology parameters are intentionally omitted.