Serverless GPU Benchmarks
Measured runs of a pinned workload per provider — time to result, serverless overhead, sample counts, and observed hardware. No synthetic scores.
Results
0 published results · filtered
No benchmark results are published yet.
How to read these numbers
Each row summarizes accepted runs for one immutable provider target, scenario, measurement origin, and observed startup mode. Matching samples accumulate across campaigns through the current publication pointer: with multiple runs we show the median, while a one-sample row is that single observation. The current workload is a pinned Parler-TTS Mini text-to-speech request.
- A cold observed scenario is configured to observe scale-from-zero behavior and retains container-session evidence. It is not a provider guarantee that every invocation was cold.
- Time to result is measured by the Cloudflare runner from the beginning of an invocation, before provider submission, until the complete result has been retrieved and validated. Current async profiles poll once per second, so nominal observation lag is up to roughly one second. Older campaigns retain their runner version and may use an earlier profile.
- Serverless overhead is time to result minus the workload's own compute phases (preprocessing, inference, postprocessing), computed per run: the time cost of the platform itself — startup or snapshot restore, queueing, and result retrieval. With a restored startup mode, application initialization happened before the request, so it is not part of the overhead.
- Requests originate from a Cloudflare Durable Object constrained to the EU with a best-effort Western Europe placement hint, not a fixed physical host.
- Compare rows only when the workload version, scenario, measurement origin, and startup mode match.