BearSignal.ai
ENGINEERING JOURNAL — BEARSIGNAL RESEARCH CORP.SYSTEM: SCANNING 10,000+ LISTED COS
[ 00 / WRITING ]

A company never lies in one place. It lies in the gaps between places — and XBRL is where those gaps live in plain sight, if you know the dialect.

When we first started ingesting structured filings at scale, we assumed XBRL was a solved problem. It is a standard, after all. Tagged data, machine-readable, one element per concept. Then we tried to compare revenue across four thousand companies and discovered that “revenue” was being reported under several hundred different tags, half of them invented by the filers themselves. The standard, it turned out, was a suggestion.

Custom-tag dialects are the first wall you hit. The taxonomy allows extensions — a company can mint its own element when the standard one doesn’t fit. Used honestly, this captures genuine nuance. Used carelessly, or deliberately, it scatters a single economic reality across a dozen idiosyncratic tags so that no two filers can be lined up without a translation layer. We built that layer the way you’d build a dictionary for a language with no grammar book: one stubborn mapping at a time, reconciling each extension back to the concept it actually represents.

Element drift is the second, and it’s quieter. A company tags an item one way this year and a slightly different way next year. Nothing is restated. Nothing is flagged. But the time series you stitch together is now comparing two different things and calling them one. We learned to treat every cross-period join as guilty until proven innocent — to check that the element under a label in 2024 is the same element that wore that label in 2022, and to leave a trail when it isn’t.

The third is the one worth losing sleep over: the silent restatement. A prior-period number changes, and the change is buried in the tagged data without a word in the narrative. No press release, no eighth-page footnote you could at least find. Just a value that used to be one thing and is now another, sitting quietly in a filing nobody reads line by line. Surfacing those — diffing what a company said about a period against what it later said about the same period — is some of the most productive reading we do.

The deeper lesson isn’t about XBRL at all. It’s that the data being public is not the same as the data being honest, or even legible. Anyone can download a filing. Reading it correctly — normalizing the dialects, anchoring the elements, keeping a record of what moved — is the unglamorous work that has to happen before any judgment is worth making.

We don’t talk much about what comes after the reading. That’s a different discipline. But none of it means anything if the foundation is a time series quietly comparing apples to a different year’s apples. Get the reading right first. Everything downstream is built on it — and so, in our shop, “read it correctly before you read anything into it” has become something close to law.