A demonstration on our showcase coverage proves very little, because we chose it. An evaluation is worth running against events you already track, where you know what happened and when you learned about it, so the comparison has a control. Bring the use case; we point the network at it.
What you are measuring is two things, and both are recomputable from the records rather than asserted by us. How much earlier a record arrived than your existing source, which the event time and the serving point of presence on each record let you calculate. And how often it was right, which is what the authenticity and data quality values are for.
The evaluation runs against live coverage with a sandbox credential, not a curated sample file, because a sample cannot demonstrate whether we are early. Research on point-in-time history runs alongside it, so a backtest and the live stream can be compared directly - they are the same records.