BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

We’ve argued that the right unit of detection is a contradiction inside a company’s own disclosures, not a statistical outlier. That commitment raises an immediate engineering question: a contradiction is a starting point, not a verdict. How do you get from “these two figures are hard to reconcile” to “this is a real problem worth a forensic investigation”?

Our answer is to refuse to collapse those two things into one step. We separate detection from judgment, and we give them different names, different mechanisms, and very different epistemic status.

The flywheel: detection without judgment

The first layer we call the flywheel. Its only job is to detect hard contradictions across a fixed set of dimensions — logical tensions internal to a single company’s own disclosures. It contains no machine learning at all. It does not score companies on a continuum of suspiciousness, it does not learn weights, it does not rank. It applies a set of explicit, auditable criteria and asks a binary question of each: is there a genuine internal contradiction here, or not?

When the flywheel finds one, it raises what we call a grey flag. The word is chosen carefully. A grey flag is not an accusation. It is a statement that a specific, explainable contradiction has been detected and is now pending confirmation. It says: something here does not add up, and it deserves a closer look. Nothing more.

Keeping this layer free of machine learning is a deliberate constraint, not a limitation we haven’t gotten around to fixing. A pure contradiction detector is fully explainable — every grey flag can be traced to the exact disclosed figures that triggered it. The moment you let a learned model into this layer, you lose that property, and you also open the door to the model learning correlations that have nothing to do with contradiction and everything to do with company size, sector, or survivorship. We keep the detection layer clean precisely so that what it produces means exactly one thing.

Ursa: judgment with reasoning

A grey flag is a question, not an answer. The second layer, which we call Ursa, is where judgment happens. Ursa takes a grey flag and the evidence around it and does something the flywheel deliberately cannot: it reasons. It builds an explicit chain — from the specific observation, to what that observation implies, to whether it points to a genuine problem, to the evidence that supports or undercuts that conclusion.

Only when Ursa’s reasoning confirms a genuine problem does a grey flag become a red flag. A red flag is a confirmed finding — the output of a reasoning process that examined the contradiction in context and concluded it is real, not benign. The two terms are not interchangeable, and treating them as if they were is, in our internal language, a serious violation. A grey flag is a detected tension awaiting confirmation. A red flag is a confirmed problem. The distance between them is the entire point of the architecture.

Why two layers instead of one

The temptation in any detection system is to build a single model that ingests everything and outputs a suspiciousness score. We’ve found that this conflates two things that should be kept apart.

Detection should be cheap, exhaustive, and explainable. You want to run it across the entire universe, catch every genuine contradiction, and be able to explain each one in plain terms. That calls for explicit criteria, not a learned model.

Judgment should be careful, contextual, and reasoned. You want it to weigh the specific situation, consider innocent explanations, and produce a chain of reasoning a human can follow and challenge. That calls for something quite different from a contradiction detector.

By separating the layers, each can be exactly what it needs to be. The flywheel never has to pretend it can render judgment; it just has to detect tensions reliably and explain them. Ursa never has to scan the whole market for raw contradictions; it inherits clean, explainable grey flags and spends its effort on the part that actually requires reasoning. The grey flag is the clean interface between the two — a detected contradiction, fully traceable, handed from a system that finds things to a system that judges them.

Most of what makes this hard is keeping the boundary honest: making sure the detection layer never quietly starts judging, and the judgment layer never quietly starts inventing contradictions the detector didn’t find. Much of our engineering discipline is, in large part, about defending that boundary.