Skip to content

Serverless GPU Pricing

Compare serverless GPU providers with per-second billing and scale-to-zero, with every price linked to its source.

Results

53 matching offerings · 83 billing options · filtered

Baseten.co
NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 0.1663/minute
    compute time
Component price — not a total
Standard
scale to zero · dedicated
1 minute increments · bills deploying, scaling up, scaling down, making predictions
Baseten.co
H100 · 80 GB VRAM/GPU
H100 · Serverless · fixed
  • USD 0.1083/minute
    compute time
Component price — not a total
Standard
scale to zero · dedicated
1 minute increments · bills deploying, scaling up, scaling down, making predictions
Baseten.co
A100 · 80 GB VRAM/GPU
A100 · Serverless · fixed
  • USD 0.06667/minute
    compute time
Component price — not a total
Standard
scale to zero · dedicated
1 minute increments · bills deploying, scaling up, scaling down, making predictions
Beam.cloud
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.001561/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Beam.cloud
1× A100 · 80 GB VRAM/GPU
A100 80GB · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.000625/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Beam.cloud
1× H100 · 80 GB VRAM/GPU
H100 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.000986/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Beam.cloud
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
RTX PRO 6000 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.000758/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Beam.cloud
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.001136/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
Cerebrium
1× A100 · 80 GB VRAM/GPU
A100 (80GB) · Serverless · fixed
  • USD 0.00000655/core second
    cpu core time
  • USD 0.000583/second
    gpu time
  • USD 0.00000222/gb second
    memory time
Component price — not a total
Standard
scale to zero
1 second increments · bills actual compute time
EU compute EU-constrained
Sweden
3 geography evidence items
Fal.ai
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 6.25/hour
    compute time
≈ EUR 5.442/hour
B200 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 6.25/hour
    compute time
≈ EUR 5.442/hour
B200 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B300 · 288 GB VRAM/GPU
B300 · Serverless · fixed
  • USD 8.5/hour
    compute time
≈ EUR 7.401/hour
B300 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA B300 · 288 GB VRAM/GPU
B300 · Serverless · fixed
  • USD 8.5/hour
    compute time
≈ EUR 7.401/hour
B300 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× H100 · 80 GB VRAM/GPU
H100 · Serverless · fixed
  • USD 3.99/hour
    compute time
≈ EUR 3.474/hour
H100 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× H100 · 80 GB VRAM/GPU
H100 · Serverless · fixed
  • USD 3.99/hour
    compute time
≈ EUR 3.474/hour
H100 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
RTX PRO 6000 · Serverless · fixed
  • USD 2.99/hour
    compute time
≈ EUR 2.603/hour
RTX PRO 6000 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
RTX PRO 6000 · Serverless · fixed
  • USD 2.99/hour
    compute time
≈ EUR 2.603/hour
RTX PRO 6000 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 4.5/hour
    compute time
≈ EUR 3.918/hour
H200 always-on
always on
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Fal.ai
1× NVIDIA H200 · 141 GB VRAM/GPU
H200 · Serverless · fixed
  • USD 4.5/hour
    compute time
≈ EUR 3.918/hour
H200 scale-to-zero
scale to zero
1 second increments · bills setup, idle, running, draining, terminating
The Serverless & Compute Pricing table publishes As low as amounts, but those discounted prices require sales contact and are omitted.
Google Cloud
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
NVIDIA RTX Pro 6000 · Serverless · fixed
  • USD 0.000018/core second
    cpu core time
  • USD 0.0003652/second
    gpu time
  • USD 0.000002/gib second
    memory time
Component price — not a total
NVIDIA RTX Pro 6000 No zonal redundancy
scale to zero
100 ms increments · 1 minute minimum · bills entire instance lifecycle
EU compute EU-constrained
Netherlands
3 geography evidence items
Google Cloud
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
NVIDIA RTX Pro 6000 · Serverless · fixed
  • USD 0.000018/core second
    cpu core time
  • USD 0.0005691/second
    gpu time
  • USD 0.000002/gib second
    memory time
Component price — not a total
NVIDIA RTX Pro 6000 Zonal redundancy
scale to zero
100 ms increments · 1 minute minimum · bills entire instance lifecycle
EU compute EU-constrained
Netherlands
3 geography evidence items
Inferless
1× A100 · 80 GB VRAM/GPU
Nvidia A100 Dedicated · Serverless · fixed
  • USD 0.001491/second
    compute time
≈ EUR 4.674/hour
Standard
scale to zero · dedicated
1 second increments · bills model loading, healthy state, request processing
Enterprise pricing is listed as discounted or custom and has no published price; omitted from the catalog.
Microsoft Azure
1× A100 · 80 GB VRAM/GPU
NVIDIA A100 · Serverless · fixed
  • USD 0.000024/core second
    cpu core time
  • USD 0.000529/second
    gpu time
  • USD 0.000003/gib second
    memory time
Component price — not a total
NVIDIA A100 — East US
scale to zero
1 second increments · bills active usage
EU compute EU-constrained
Italy · Sweden
5 geography evidence items
Microsoft Azure
1× A100 · 80 GB VRAM/GPU
NVIDIA A100 · Serverless · fixed
  • USD 0.000024/core second
    cpu core time
  • USD 0.000688/second
    gpu time
  • USD 0.000003/gib second
    memory time
Component price — not a total
NVIDIA A100 — Sweden Central
scale to zero
1 second increments · bills active usage
EU compute EU-constrained
Italy · Sweden
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00001965/core second effective
    cpu core time
  • USD 0.000694/second → USD 0.001041/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.00000333/gib second effective
    memory time
Component price — not a total
Broad region
scale to zero
none minimum · bills actual compute time
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00002293/core second effective
    cpu core time
  • USD 0.000694/second → USD 0.001215/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.000003885/gib second effective
    memory time
Component price — not a total
Narrow region
scale to zero
none minimum · bills actual compute time
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second
    cpu core time
  • USD 0.000694/second
    gpu time
  • USD 0.00000222/gib second
    memory time
Component price — not a total
Standard region
scale to zero
none minimum · bills actual compute time
5 geography evidence items
OVHcloud
4× NVIDIA H200 · 131.25 GB VRAM/GPU
h200-4-gpu · Serverless · fixed
  • USD 27.28/hour
    compute time
≈ EUR 23.75/hour
h200-4-gpu always-on
always on · dedicated
1 minute increments · bills SCALING, RUNNING
OVHcloud
4× NVIDIA H200 · 131.25 GB VRAM/GPU
h200-4-gpu · Serverless · fixed
  • USD 27.28/hour
    compute time
≈ EUR 23.75/hour
h200-4-gpu scale-to-zero
scale to zero · dedicated
1 minute increments · bills SCALING, RUNNING
OVHcloud
1× H100 · 80 GB VRAM/GPU
h100-1-gpu · Serverless · fixed
  • USD 3.39/hour
    compute time
≈ EUR 2.952/hour
h100-1-gpu always-on
always on · dedicated
1 minute increments · bills SCALING, RUNNING
OVHcloud
1× H100 · 80 GB VRAM/GPU
h100-1-gpu · Serverless · fixed
  • USD 3.39/hour
    compute time
≈ EUR 2.952/hour
h100-1-gpu scale-to-zero
scale to zero · dedicated
1 minute increments · bills SCALING, RUNNING
OVHcloud
1× A100 · 80 GB VRAM/GPU
a100-1-gpu · Serverless · fixed
  • USD 3.35/hour
    compute time
≈ EUR 2.917/hour
a100-1-gpu always-on
always on · dedicated
1 minute increments · bills SCALING, RUNNING
OVHcloud
1× A100 · 80 GB VRAM/GPU
a100-1-gpu · Serverless · fixed
  • USD 3.35/hour
    compute time
≈ EUR 2.917/hour
a100-1-gpu scale-to-zero
scale to zero · dedicated
1 minute increments · bills SCALING, RUNNING
Replicate
8× A100 · 80 GB VRAM/GPU
8x Nvidia A100 (80GB) GPU · Serverless · fixed
  • USD 0.0112/second
    compute time
≈ EUR 35.11/hour
Committed multi-GPU capacity
request metered · commitment
bills request execution time
Replicate
4× A100 · 80 GB VRAM/GPU
4x Nvidia A100 (80GB) GPU · Serverless · fixed
  • USD 0.0056/second
    compute time
≈ EUR 17.55/hour
Committed multi-GPU capacity
request metered · commitment
bills request execution time
Replicate
4× H100 · 80 GB VRAM/GPU
4x Nvidia H100 GPU · Serverless · fixed
  • USD 0.0061/second
    compute time
≈ EUR 19.12/hour
Committed multi-GPU capacity
request metered · commitment
bills request execution time
Replicate
2× H100 · 80 GB VRAM/GPU
2x Nvidia H100 GPU · Serverless · fixed
  • USD 0.00305/second
    compute time
≈ EUR 9.56/hour
Committed multi-GPU capacity
request metered · commitment
bills request execution time
Replicate
8× H100 · 80 GB VRAM/GPU
8x Nvidia H100 GPU · Serverless · fixed
  • USD 0.0122/second
    compute time
≈ EUR 38.24/hour
Committed multi-GPU capacity
request metered · commitment
bills request execution time
Runpod
1× NVIDIA RTX PRO 6000 Blackwell · 96 GB VRAM/GPU
RTX 6000 Pro · Serverless · fixed
  • USD 3.49/hour
    gpu time
≈ EUR 3.039/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active worker discounts are available through sales inquiry and are not published; custom pricing plans are available through sales inquiry, so no Active or custom-priced variants were included.
Runpod offers custom pricing plans for large scale and enterprise workloads; contact sales for pricing.
Runpod
1× A100 · 80 GB VRAM/GPU
A100 · Serverless · fixed
  • USD 2.72/hour
    gpu time
≈ EUR 2.368/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active worker discounts are available through sales inquiry and are not published; custom pricing plans are available through sales inquiry, so no Active or custom-priced variants were included.
Runpod offers custom pricing plans for large scale and enterprise workloads; contact sales for pricing.
Runpod
1× NVIDIA B200 · 180 GB VRAM/GPU
B200 · Serverless · fixed
  • USD 8.64/hour
    gpu time
≈ EUR 7.523/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active worker discounts are available through sales inquiry and are not published; custom pricing plans are available through sales inquiry, so no Active or custom-priced variants were included.
Runpod offers custom pricing plans for large scale and enterprise workloads; contact sales for pricing.