BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

The Record Is the Product

[ /methodology ] · 2026-07-15 · 4 min read

A red flag on this platform is not an opinion. It is the output of a trained machine — and the thing we actually sell is the record of how that machine learned to make it.

That sentence took us a while to get right, and the earlier versions were wrong in a way worth being precise about.

Who does what

Ursa, our reasoning engine, reads filings and issues flags. No analyst opinion sits behind an individual flag. No committee votes on it. A flag is what the engine produces when the evidence in a company’s own disclosures crosses the line the engine has learned.

Annotation experts — credentialed, practising forensic accountants — do something categorically different. They work historical filings, writing signed, blind reasoning chains about cases that are already closed. Their annotations are what the engine learns from. They label the past; they issue no opinion on any live company.

Both halves matter. Drop the first and you are describing a research firm with a fast database. Drop the second and you are describing a model with no traceable source of judgment. The product is the pair: a machine that issues the finding, and a record showing exactly whose judgment taught it what a finding looks like.

Why the boundary is not cosmetic

The earliest version of our public description said something close to “three accountants sign every flag.” It read well. It was also wrong in a way that mattered, and we took it down when we saw it.

Wrong descriptively: no accountant signs a live flag. The engine issues it. Wrong in a heavier sense too — a description in which credentialed accountants put their names to findings about currently listed companies claims something about the nature of the work, and about what those professionals are doing, that is not what happens. The annotators are teaching. They are not issuing opinions on issuers.

Getting that wrong in our own copy was not a wording slip. It came from describing an aspiration — the feel of expert-backed judgment — instead of the mechanism. The correction was to write down the mechanism and let it read how it reads.

Why a record, rather than a score

The industry default is a number. A risk score, a percentile, a letter grade. It compresses well and it is easy to consume.

We think it destroys the only thing worth buying. A score tells you the system’s conclusion and nothing about its basis, so the only available response is to trust it or not. A record tells you which dimensions of the company’s own filings contributed evidence, what the historical judgments behind that pattern were, who made them, where they disagreed, and how the disagreement was resolved. You can argue with the second. You cannot argue with the first — you can only accept or discard it.

This produces obligations we would not otherwise have.

Reasoning chains are blindthree readers who cannot see each other’s work, because agreement that was copied is not agreement.

Disagreements are published, not resolved away. A record containing only unanimous cases has been filtered for the easy ones, and the filter is invisible to the reader.

Nothing is silently rewritten. Corrections supersede; originals stay. Our own ledger enforced that against us before any customer tested it.

And the corpus is guarded. Machine-generated reasoning never enters it. An enforcement database is a textbook, never an answer key — train on cases that got caught and you learn the shape of getting caught.

What Ursa is, and is not

Ursa is not a judge, and describing it as one would be the same class of error as the sentence we withdrew.

It is a reasoning engine that has learned from a signed corpus and issues gray flags — detected hard contradictions, pending confirmation. Pending is doing real work in that phrase. A flag says the numbers in a filing cannot all be true and states which dimensions carry that evidence. It does not say fraud, and the distance between those is the distance between research and an accusation.

Where we intend to end up is a machine that renders forensic judgment at market scale. Where we are is a machine that surfaces contradictions and a growing record of expert reasoning teaching it what those contradictions mean. Saying the second while meaning the first is how a company ends up with copy it has to take down.

The uncomfortable part

The record is only as good as the judgment in it, and judgment does not scale by hiring faster. Every annotation is hours of an experienced professional’s attention on a closed case that will teach the engine one thing. There is no shortcut — the shortcut is precisely the machine-generated reasoning we refuse.

So the honest description of our position: the engine’s ceiling is set by the corpus, the corpus is set by the people writing it, and that is the constraint we are working against rather than around. We would rather be limited by something real than fast on something hollow.


All figures are system-level results current as of publication date; methodology parameters are intentionally omitted.