BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

Two Kinds of Evidence, and Only One of Them Can Raise a Flag

[ /engineering ] · 2026-07-16 · 5 min read

We have written before about why we look for contradictions rather than outliers. That piece made the case in the abstract. This one is about what happened when we had to turn it into a schema — because the moment you build the thing, you discover that the clean distinction has a messy edge, and the edge is where the design lives.

Here is the problem in one sentence: we do use relative measures. We could not detect anything useful without them. So a rule that says “only absolute contradictions count” is either a lie about our own system, or it throws away half the evidence. Neither is acceptable. What we ended up with is a two-class taxonomy where both kinds of evidence are collected, both are shown, and only one of them is allowed to raise a flag on its own.

The two classes

The first class is internal contradiction. These are the dimensions where a company’s own reported numbers cannot all be true at once, without any reference to what other companies look like. Profit that never becomes cash. Receivables growing faster than the revenue that supposedly created them. Inventory growing faster than the sales it is supposedly there to serve. Cash on the balance sheet alongside borrowing that costs real interest. Interest expense that does not reconcile with the cash the company says it holds.

You do not need a peer group to evaluate any of those. You need the company’s own filing, and arithmetic. If receivables outrun revenue for long enough, either the revenue is not being collected or it was never really there — and both of those are things a reader is entitled to ask about. The claim is about the document, not about the industry.

The second class is level. Soft-asset intensity. Gross margin. These say: this company sits at an unusual point relative to companies like it. That is real information. A gross margin far from what comparable operators achieve is worth knowing. But the statement is inherently about the comparison set, and that is exactly the weakness — change the peer group and the reading changes with it.

The rule that came out of it

The gate we settled on is structural, not statistical: a flag requires at least one internal contradiction to have fired. Level dimensions can add to a case, sharpen it, raise its tier. They cannot start one. A company whose only distinguishing feature is that its margin looks unusual for its sector does not get flagged, no matter how unusual.

This is not a preference for one kind of math over another. It is about what we can defend when someone pushes back. If our entire case is “your margin is unlike your peers,” the answer “our business is not like theirs” is not a dodge — it may well be correct, and we have no way to distinguish. If the case is “your receivables have outrun your revenue for several consecutive periods,” the response has to engage with the company’s own numbers. One of those conversations goes somewhere. The other one ends.

There is a second-order effect that mattered more than we expected. Level dimensions need a comparison group, and choosing that group is a modelling decision with a lot of freedom in it. Sector? Size band? Both? How coarse? Every one of those choices moves the answer. When a level dimension can trigger on its own, all that discretion flows straight into the output, and you end up with a detector whose sensitivity is a function of how you drew the buckets. Demoting them to supporting evidence means that discretion can only ever adjust a case that already has a document-level basis. The freedom is still there; it just cannot be the whole story.

The one that had to be handled separately

Tax divergence — where a company’s tax expense does not track the profit it reports — sits awkwardly between the classes. It has the shape of an internal contradiction: two numbers in the same filing that ought to move together and do not. But it has a long list of entirely legitimate explanations. Carried-forward losses. A one-off gain taxed differently. A genuine incentive regime. In a market where those are ordinary, a rule that treats tax divergence as a hard contradiction will spend most of its time pointing at companies that have done nothing wrong.

So it gets its own handling: on its own it can only ever produce our lowest tier — a note, not a flag. Combined with a genuine internal contradiction, it counts fully. This is the shape of a compromise, and we would rather write it down as one than pretend the taxonomy came out clean.

What it looks like from the outside

None of this is visible in the output as a hierarchy. A published gray flag lists the dimensions that fired, and a reader can see for themselves which are document-internal and which are comparative. We are not asking anyone to trust that we weighted them correctly. We are showing which ones spoke, and the rule about what is allowed to raise a flag is written down rather than buried in a coefficient.

That last point is the part we would defend hardest. A weighting scheme can encode the same rule — you could give level dimensions a small enough weight that they never cross the line alone. It would behave almost identically and it would be much worse, because the day someone retunes the weights, the rule quietly stops existing and nothing announces it. A structural gate fails loudly. A small coefficient fails silently, and by the time you notice, you have been publishing on a different standard for months.

We have made the coefficient mistake before, in a different part of the system, and the thing that eventually caught it was not a test. It was someone reading the output and asking why a particular company was on the list.


All figures are system-level results current as of publication date; methodology parameters are intentionally omitted.