BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

Evidence That Survives Being Watched

[ /engineering ] · 2026-07-18 · 5 min read

A scanner has a property that breaks classical statistics, and it breaks it quietly enough that you can ship for a year before noticing.

The property is this: the scanner never stops looking. A filing arrives, we score it. Another arrives a week later, we score again. A restatement lands, a new period closes, an exchange announcement drops — every one of those is another look at the same company. Over a year a single name gets examined dozens of times, and we would like to be able to say, at any moment, “the evidence against this company’s numbers is now strong enough to act on.”

If you build that on p-values, you have built something that lies.

Why repeated looks are fatal

A p-value is calibrated for one look. You fix the test, you collect the data, you compute the number once. The guarantee — that a false positive happens at most five percent of the time — is a statement about that single act.

Peek repeatedly and the guarantee dissolves. Test the same company after every new filing, and stop the moment the number crosses your line, and you will cross that line eventually on a perfectly ordinary company, purely because you kept asking. This is not a subtle effect. It is the reason clinical trials pre-register their interim analyses and pay a penalty for each one. A scanner that re-tests thousands of companies continuously is doing the pathological version of this, at scale, forever.

The standard fixes do not fit. Multiple-testing corrections need to know how many tests you will run — but a scanner has no last test; it runs until the company delists. Pre-registered interim looks need a schedule — but filings arrive when they arrive. Every classical remedy assumes an experiment with an end, and we do not have one.

The other kind of evidence measure

The construct we use instead is an e-value. Its definition is almost aggressively simple: a non-negative quantity whose expected value, if the null hypothesis is true, is at most one.

That is the whole thing. Under “nothing is wrong here,” the number averages to one or less. So a large value is, directly and interpretably, evidence against the null — not a probability, not a score, just a statement of how much the observation disfavours the boring explanation.

Two consequences follow, and both are the reason we built on it.

It multiplies. If two pieces of evidence are independent, the expectation of their product is the product of their expectations, which is still at most one. The product of e-values is itself an e-value. And under dependence there is a stronger construction: a sequence where each term is built conditional on everything seen before produces a running product that remains valid regardless of how the pieces relate. That running product is a test martingale, and it is the object our accumulating evidence actually is.

It survives being watched. Ville’s inequality says the probability that such a process ever reaches a given multiple of one is bounded by its reciprocal. Ever. Not “at the point you planned to stop” — at any stopping time, including one you chose after looking. This is what “anytime-valid” means, and it is precisely the property a scanner needs. New filing arrives, multiply it in, read the number. The false-positive guarantee is intact. No penalty, no correction, no recomputation of the whole universe.

What changes when you build on it

The first thing that changes is the null hypothesis, and this turned out to matter more than the machinery.

A likelihood-ratio approach — the obvious sequential-Bayes construction — needs a probability of the evidence given a bad company. That requires a definition of “bad,” which in practice means a labelled set of known cases. We have written elsewhere about why we refuse to let an enforcement database define what fraud looks like; the short version is that it teaches the system to find the kind of fraud that got caught. But there is a plainer engineering objection: that anchor makes the whole detection layer dependent on a label set we do not control and cannot audit.

The e-value construction needs no such thing. The null is “this company is an ordinary company — each of its features is a typical draw from what the market looks like.” Ordinary is definable from the data itself. Nothing labelled enters. What accumulates is evidence against ordinariness, which is a claim we can actually support.

The second thing that changes is what a missing value does. Under a scoring scheme, an absent input has to be imputed or filled with a neutral value, and both are quiet fabrications. Under this one it is exact: no data means no bet, and no bet is a factor of one. It multiplies in and changes nothing. The arithmetic identity of “one” is the mathematical form of the rule we already held on principle — that we do not invent numbers we do not have.

The part that is genuinely hard

None of the above rescues you from dependence, and dependence is the real problem in this domain.

Financial features are not independent. Several accrual measures are near-restatements of one another; several distress models share inputs. Multiply a family of correlated signals together as if they were separate witnesses and the running product inflates fast — not because the company is unusual, but because you counted one observation five times. The mathematics is still valid; the input assumption is not.

That is not a footnote to this design. It is the part that required actual work, and it deserves its own post rather than a paragraph here.

What is worth saying now is the shape of the tradeoff. The classical apparatus gives you sharp guarantees on the condition that you look once, which you cannot honour. This one gives you guarantees that hold under continuous observation, on the condition that you are honest about how your evidence relates. We would rather have the assumption we can inspect than the one we are structurally guaranteed to violate.


All figures are system-level results current as of publication date; methodology parameters are intentionally omitted.