Skip to content

Serverless GPU Pricing

Compare serverless GPU providers with per-second billing and scale-to-zero, with every price linked to its source.

Results

13 matching offerings · 21 billing options · filtered

Baseten.co
NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 0.1663/minute
    compute time
Component price — not a total
Standard
scale to zero · dedicated
1 minute increments · bills deploying, scaling up, scaling down, making predictions
Beam.cloud
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.001561/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Beam.cloud
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.001136/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Fal.ai
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 6.25/hour
    compute time
≈ EUR 5.442/hour
B200 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 6.25/hour
    compute time
≈ EUR 5.442/hour
B200 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B300 · 288 GB VRAM/GPU
B300 · Serverless · fixed
  • USD 8.5/hour
    compute time
≈ EUR 7.401/hour
B300 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B300 · 288 GB VRAM/GPU
B300 · Serverless · fixed
  • USD 8.5/hour
    compute time
≈ EUR 7.401/hour
B300 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 4.5/hour
    compute time
≈ EUR 3.918/hour
H200 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 4.5/hour
    compute time
≈ EUR 3.918/hour
H200 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Runpod
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 8.64/hour
    gpu time
≈ EUR 7.523/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active worker discounts are available through sales inquiry and are not published; custom pricing plans are available through sales inquiry, so no Active or custom-priced variants were included.
Runpod offers custom pricing plans for large scale and enterprise workloads; contact sales for pricing.
Runpod
1× NVIDIA B300 · 280 GB VRAM/GPU
B300 · Serverless · fixed
  • USD 9.98/hour
    gpu time
≈ EUR 8.69/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active worker discounts are available through sales inquiry and are not published; custom pricing plans are available through sales inquiry, so no Active or custom-priced variants were included.
Runpod offers custom pricing plans for large scale and enterprise workloads; contact sales for pricing.