Benchmarks
Measured against going direct
A gateway adds a network hop. The honest question is how much that costs you, so every benchmark here compares Gatewayz with calling the same provider directly.
No benchmark run published yet
This page shows measured results only. Until a run against production meets the rules below, there are no latency or cost figures to show, and we won't estimate them.
Method
- The same set of tasks is sent through Gatewayz and, where we hold that provider's key, straight to the provider.
- Each task runs several times per route; we take the median for each task, then across tasks.
- We record end-to-end latency and, where the stream reports it, time to first token.
Rules for what gets published
These rules are enforced by the benchmark harness itself, not applied by hand afterwards.
- No baseline, no number. A latency figure is published only next to a direct-to-provider measurement from the same run.
- Within 15% or not at all. Gateway latency is published only when its median is within 15% of the fastest direct route. If we're slower than that, the page says so rather than showing a flattering subset.
- No cost claim without cache reads. A cost comparison is published only when the run actually read from the prompt cache and has a direct baseline. If a provider doesn't report cache usage, it counts as unmeasured, not as zero.
- Nothing fabricated. If a stream doesn't report time to first token, that column stays empty.
Further reading
The reasoning behind these rules: how to benchmark a gateway honestly.
Measure it yourself
Point the same workload at Gatewayz and at the provider, and compare.
