Demo

The real-time data layer for markets and machines

The events that move markets and the data that trains models happen in the real world first. Why a real-time layer is becoming its own category.

Every interesting signal starts as an event in the physical or digital world: a fabrication plant changes utilization, a grid interconnect filing appears, a satellite passes overhead, or a launch window opens. By the time that event reaches a clean, packaged vendor dataset, most of its value has already been priced in or has gone stale for training. The gap between when something happens and when you can act on it - the latency nobody puts on a datasheet - is where the edge lives.

The gap is not uniform, either, which is what makes it exploitable rather than merely annoying. Some events are public within seconds and simply unnoticed; others are observable in the physical world long before anyone writes them down. A vendor that only reads what has already been published inherits the second kind late by construction, no matter how fast its pipeline is - which is why we describe the network by where it listens before we describe what it does with what it hears.

Closing that gap is an infrastructure problem, not a dashboard. It means edge nodes listening to real-world events worldwide, normalizing and timestamping each observation at the receiving node, evaluating the raw stream into signal rather than passing through unfiltered data, and delivering it - by stream or in bulk - the moment it happens. Each of those steps is hard on its own; doing all of them together, reliably, at low latency, is the product.

We think about it as one layer with two customers. Hedge funds want real-time, ground-truth signal to trade on. AI teams want large, fresh, high-signal training data to build on. The same listening and evaluation pipeline serves both - markets and machines - which is why we build it as shared infrastructure rather than two separate products.

Calling it a layer is a claim about where the boundary sits, and it is worth being precise. A layer is something other systems are built on without reimplementing it: nobody writes their own TLS stack to make an HTTPS request, and nobody should have to run listening infrastructure in nineteen regions to know that a grid interconnect filing appeared. Today most firms do exactly that, because the alternative on offer is a packaged dataset that arrives too late to be an input to anything live. The work is duplicated across every buyer, badly, and it is nobody's core business.

What has changed is the composition of what moves capital. The events that reprice a sector increasingly happen in physical infrastructure - compute, chips, power, minerals, orbit, and the networks between them - rather than in a quarterly disclosure about them. That is a harder collection problem than parsing filings, and it is the reason the layer has to start at the point of observation rather than at a vendor's export. You cannot be earlier than the source you resell.

The same shift explains why the second buyer appeared at all. A model grounded in a corpus assembled last year describes a world that has moved, and the fix is not a larger corpus but a current one, with provenance good enough to cite. Grounding turned out to need the same properties a research desk has always demanded: resolved entities, honest timestamps, and evidence attached to the record rather than filed separately.

The rest of this blog goes deeper on the pieces, and mostly on the parts that are harder than they look: why polling is late before anything breaks, why a quiet source is an outage, why the clock is a correctness requirement rather than plumbing, and why we label the noise instead of deleting it. The glossary is where the vocabulary those posts assume is defined.

Share