BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

Forty Days of Looking Healthy

[ /engineering ] · 2026-08-04 · 5 min read

Our engine kept marking companies. Our scanner kept ingesting filings. Our public case count stayed comfortably non-zero and the site returned 200 to every check we had.

For forty days, nothing new reached the public surface.

Every monitor we owned was green throughout, and each one was telling the truth. This post is about why that combination is possible, and what we changed so that it is not.

Two failures, stacked

The first: the publisher had no schedule.

Every other stage in the pipeline ran on a timer. The publisher — the step that carries an engine judgment onto the public surface — was invoked by hand. It had always been invoked by hand, back when a person was watching the pipeline daily and the manual step was indistinguishable from an automatic one. Then attention moved elsewhere, and a stage with no schedule stopped, and a stage with no schedule looks exactly the same stopped as running.

The second: an assembly step quietly starved.

Before publication, each case is assembled into its outward form, including the list of dimensions that fired. There is a hard gate on that list: if we cannot articulate why a company was flagged, it does not get published. No explanation, no publication. That gate is correct and we would build it again.

But the assembly layer’s vocabulary had been written for one market. Applied to companies from the other, it could not translate the engine’s dimension names into anything it recognised — so it produced an empty explanation. Not an error. An empty list, which the gate then correctly refused to publish.

The gate did its job perfectly. It was fed a lie by the layer upstream, and the lie was “this case has no articulable basis,” when the truth was “this layer does not speak this market’s language.”

Why the monitors could not see it

Each monitor asked about one stage. Is ingestion running? Yes. Is the engine marking? Yes. Does the public surface have cases? Yes — hundreds, all published before the stall. Is the site up? Yes.

Nothing asked whether the newest engine output had reached the surface. The pipeline was monitored as a set of stations, and what failed was the track between two of them. A stalled conveyor between two healthy stations produces exactly the readings we were getting.

The stock level made it worse. Because the published count was large and non-zero, every count-based check was reassuring. A non-zero inventory is not evidence that anything is being produced. It took a stall of over a month before any age-based measure would have looked odd, and we had no age-based measure.

What we changed

Lag between stages, not health of stages. The chain now carries checks that compare timestamps across the boundary — the newest engine mark against the newest publication. If judgments stop reaching the surface, that gap grows immediately, regardless of how healthy either end looks. This is the only check in the set that could have caught the original failure, and it did not exist.

A stage with no scheduler is a red condition, not a blank. Three stages had no scheduler at all, and rendered as unknown. Unknown reads as “probably fine” to a human scanning a board. They now render red with the reason stated, because a stage nobody schedules is a stage that will stop and not tell you.

The lamps read the scheduler’s state, not just data freshness. Disable a schedule and freshness stays fine for days before it degrades. Reading the scheduler means the board goes red the moment it is switched off, rather than the week after.

And we test the lamps by breaking things. A drill disables each schedule in turn, confirms the corresponding lamp turns red, restores it, and confirms it turns green. Self-tests only prove the judgment functions are correct; they say nothing about whether the wiring reaches the real infrastructure. A monitor that has never been observed to go red is a monitor with an untested claim — the same acceptance standard we apply to constraints.

The thirty-one

There was a second cost, discovered while tracing the first. A cleanup pass months earlier had removed cases whose explanation field was empty — reasonable at the time, since an empty explanation meant an unpublishable case.

But some of those were empty for the reason above: the assembly layer could not translate them, not because there was nothing to say. Thirty-one cases had been withdrawn from the public surface for a defect in a translation layer that had nothing to do with their merits.

When the translation was fixed they were restored — the original records, reinstated, not re-created as new ones. That distinction mattered enough to do the harder version: their original identity and history stay intact, so the record shows a withdrawal and a reinstatement rather than a fresh case that happens to look the same.

The lesson we keep relearning

Every component was honest. Ingestion honestly reported ingestion. The gate honestly refused a case with no stated basis. The site honestly returned 200.

Nobody was responsible for the question “did the previous stage’s output reach me?” — and that question is where systems of honest components go wrong. It is not a monitoring gap in the usual sense. It is a gap in whose job it is to notice, and no amount of per-component rigour fills it.


All figures are system-level results current as of publication date; methodology parameters are intentionally omitted.