Skip to content

61 offerings from 17 providers across serverless products and GPU instances.

80 complete comparable prices; incomplete and additive prices remain visible but are not ranked as totals.

21 of 25 serverless offerings include a scale-to-zero operating mode.

Published configurations: 40 GB · 80 GB

Serverless prices

All serverless GPUs →
Microsoft Azure
1× A100 · 80 GB VRAM/GPU
NVIDIA A100 · Serverless · fixed
  • USD 0.000024/core second
    cpu core time
  • USD 0.000529/second
    gpu time
  • USD 0.000003/gib second
    memory time
Component price — not a total
eastus
scale to zero
1 second increments · bills active
EU compute EU-constrained
United States · Italy · Sweden
stale Verified 17 Sept 2026 Official price source (opens in new tab)
6 geography evidence items
Microsoft Azure
1× A100 · 80 GB VRAM/GPU
NVIDIA A100 · Serverless · fixed
  • USD 0.000024/core second
    cpu core time
  • USD 0.000688/second
    gpu time
  • USD 0.000003/gib second
    memory time
Component price — not a total
swedencentral
scale to zero
1 second increments · bills active
EU compute EU-constrained
United States · Italy · Sweden
stale Verified 17 Sept 2026 Official price source (opens in new tab)
6 geography evidence items
Baseten.co
A100 · 80 GB VRAM/GPU
A100 · Serverless · fixed
  • USD 0.06667/minute
    compute time
Component price — not a total
A100 standard
scale to zero · dedicated
1 minute increments · bills deploying, scaling up, scaling down, making predictions
Beam.cloud
1× A100 · 80 GB VRAM/GPU
A100 80GB · Serverless · fixed
  • USD 0.0000125/core second
    cpu core time
  • USD 0.000625/second
    gpu time
  • USD 0.0000021/gib second
    memory time
Component price — not a total
Standard
scale to zero
1 ms increments · bills code running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
The Serverless committed-spend option is presented as “With committed spend Book a call →” without a published amount; it was omitted.
The Serverless section includes a “With committed spend” / “Book a call” option without a published amount; it was omitted.
Cerebrium
1× A100 · 40 GB VRAM/GPU
A100 (40GB) · Serverless · fixed
  • USD 0.00000655/core second
    cpu core time
  • USD 0.000555/second
    gpu time
  • USD 0.00000222/gb second
    memory time
Component price — not a total
Standard scale-to-zero
scale to zero
1 second increments
Enterprise pricing is listed as Custom and was omitted.
1 geography evidence item
Cerebrium
1× A100 · 80 GB VRAM/GPU
A100 (80GB) · Serverless · fixed
  • USD 0.00000655/core second
    cpu core time
  • USD 0.000583/second
    gpu time
  • USD 0.00000222/gb second
    memory time
Component price — not a total
Standard scale-to-zero
scale to zero
1 second increments
Enterprise pricing is listed as Custom and was omitted.
1 geography evidence item
Cerebrium
1× A100 · 40 GB VRAM/GPU
A100 (40GB) · Serverless · fixed
  • USD 0.00000655/core second
    cpu core time
  • USD 0.000555/second
    gpu time
  • USD 0.00000222/gb second
    memory time
Component price — not a total
Standard
scale to zero
1 second increments · bills active compute time
Enterprise pricing is listed as Custom and was omitted.
1 geography evidence item
Cerebrium
1× A100 · 80 GB VRAM/GPU
A100 (80GB) · Serverless · fixed
  • USD 0.00000655/core second
    cpu core time
  • USD 0.000583/second
    gpu time
  • USD 0.00000222/gb second
    memory time
Component price — not a total
Standard
scale to zero
1 second increments · bills active compute time
Enterprise pricing is listed as Custom and was omitted.
1 geography evidence item
Inferless
1× A100 · 80 GB VRAM/GPU
Nvidia A100 Dedicated · Serverless · fixed
  • USD 0.001491/second
    compute time
≈ EUR 4.727/hour
Standard
scale to zero · dedicated
1 second increments · bills model loading, healthy state, request processing
Inferless
1× A100 · 40 GB VRAM/GPU
Nvidia A100 Shared · Serverless · fixed
  • USD 0.000745/second
    compute time
≈ EUR 2.362/hour
Standard
scale to zero
1 second increments · bills model loading, healthy state, request processing
Modal
1× A100 · 40 GB VRAM/GPU
Nvidia A100, 40 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00001507/core second effective
    cpu core time
  • USD 0.000583/second → USD 0.0006705/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.000002553/gib second effective
    memory time
Component price — not a total
Broad region
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
Modal
1× A100 · 40 GB VRAM/GPU
Nvidia A100, 40 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00002293/core second effective
    cpu core time
  • USD 0.000583/second → USD 0.00102/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.000003885/gib second effective
    memory time
Component price — not a total
Narrow region
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
Modal
1× A100 · 40 GB VRAM/GPU
Nvidia A100, 40 GB · Serverless · fixed
  • USD 0.0000131/core second
    cpu core time
  • USD 0.000583/second
    gpu time
  • USD 0.00000222/gib second
    memory time
Component price — not a total
Standard
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00001507/core second effective
    cpu core time
  • USD 0.000694/second → USD 0.0007981/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.000002553/gib second effective
    memory time
Component price — not a total
Broad region
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second → USD 0.00002293/core second effective
    cpu core time
  • USD 0.000694/second → USD 0.001215/second effective
    gpu time
  • USD 0.00000222/gib second → USD 0.000003885/gib second effective
    memory time
Component price — not a total
Narrow region
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
Modal
1× A100 · 80 GB VRAM/GPU
Nvidia A100, 80 GB · Serverless · fixed
  • USD 0.0000131/core second
    cpu core time
  • USD 0.000694/second
    gpu time
  • USD 0.00000222/gib second
    memory time
Component price — not a total
Standard
scale to zero
none minimum · bills running
stale Verified 17 Sept 2026 Official price source (opens in new tab)
5 geography evidence items
OVHcloud
1× A100 · 80 GB VRAM/GPU
a100-1-gpu · Serverless · fixed
  • USD 3.35/hour
    compute time
≈ EUR 2.95/hour
scale-to-zero
scale to zero · dedicated
1 minute increments · bills SCALING, RUNNING
Replicate
2× A100 · 80 GB VRAM/GPU
2x Nvidia A100 (80GB) GPU · Serverless · fixed
  • USD 0.0028/second
    compute time
≈ EUR 8.877/hour
Public models
request metered
bills request execution time
stale Verified 17 Sept 2026 Official price source (opens in new tab)
Replicate
4× A100 · 80 GB VRAM/GPU
4x Nvidia A100 (80GB) GPU · Serverless · fixed
  • USD 0.0056/second
    compute time
≈ EUR 17.75/hour
Public models
request metered · commitment
bills request execution time
stale Verified 17 Sept 2026 Official price source (opens in new tab)
Replicate
8× A100 · 80 GB VRAM/GPU
8x Nvidia A100 (80GB) GPU · Serverless · fixed
  • USD 0.0112/second
    compute time
≈ EUR 35.51/hour
Public models
request metered · commitment
bills request execution time
stale Verified 17 Sept 2026 Official price source (opens in new tab)
Runpod
1× A100 · 80 GB VRAM/GPU
A100 · Serverless · fixed
  • USD 2.72/hour
    gpu time
≈ EUR 2.395/hour
Flex
scale to zero
1 second increments · bills initialization, execution, idle timeout
Active workers have discounts available through sales inquiry; no public price is published, so Active variants were omitted.

Instance prices

All GPU instances →
Verda
1× A100 SXM4 · 40 GB/GPU
1× A100
40 GB VRAM/GPU · 22 vCPUs · 120 GB RAM · storage unknown
Region: unknown
  • USD 1.29/hour
    compute time
≈ EUR 1.136/hour
On-demand
standard
Verda
1× A100 SXM4 · 80 GB/GPU
1× A100
80 GB VRAM/GPU · 22 vCPUs · 120 GB RAM · storage unknown
Region: unknown
  • USD 1.79/hour
    compute time
≈ EUR 1.576/hour
On-demand
standard
Microsoft Azure
Standard_NC24ads_A100_v4
1× A100
80 GB VRAM/GPU · 24 vCPUs · 220 GB RAM · 1024 GB storage
Region: East US (Virginia) · US · eastus
  • USD 3.673/hour
    compute time
≈ EUR 3.235/hour
EASTUS Linux pay-as-you-go
standard
Google Cloud
a2-highgpu-1g
1× A100
40 GB VRAM/GPU · 12 vCPUs · 85 GB RAM · storage unknown
Region: Council Bluffs, Iowa · US · us-central1
  • USD 3.673/hour
    compute time
≈ EUR 3.235/hour
US-CENTRAL1 on-demand
standard
Google Cloud
a2-highgpu-1g
1× A100
40 GB VRAM/GPU · 12 vCPUs · 85 GB RAM · storage unknown
Region: Eemshaven, Netherlands · NL · europe-west4 · EU
  • USD 3.748/hour
    compute time
≈ EUR 3.301/hour
EUROPE-WEST4 on-demand
standard
Lambda
2× NVIDIA A100 PCIe
2× A100
40 GB VRAM/GPU · 60 vCPUs · 450 GB RAM · 1024 GB storage
Region: unknown
  • USD 1.99/hour → USD 3.98/hour effective
    compute time
≈ EUR 3.505/hour
On-demand
standard
Microsoft Azure
Standard_NC24ads_A100_v4
1× A100
80 GB VRAM/GPU · 24 vCPUs · 220 GB RAM · 1024 GB storage
Region: West Europe (Netherlands) · NL · westeurope · EU
  • USD 4.775/hour
    compute time
≈ EUR 4.205/hour
WESTEUROPE Linux pay-as-you-go
standard
Google Cloud
a2-highgpu-2g
2× A100
40 GB VRAM/GPU · 24 vCPUs · 170 GB RAM · storage unknown
Region: Council Bluffs, Iowa · US · us-central1
  • USD 7.347/hour
    compute time
≈ EUR 6.47/hour
US-CENTRAL1 on-demand
standard
Google Cloud
a2-highgpu-2g
2× A100
40 GB VRAM/GPU · 24 vCPUs · 170 GB RAM · storage unknown
Region: Eemshaven, Netherlands · NL · europe-west4 · EU
  • USD 7.496/hour
    compute time
≈ EUR 6.601/hour
EUROPE-WEST4 on-demand
standard
Lambda
4× NVIDIA A100 PCIe
4× A100
40 GB VRAM/GPU · 120 vCPUs · 900 GB RAM · 1024 GB storage
Region: unknown
  • USD 1.99/hour → USD 7.96/hour effective
    compute time
≈ EUR 7.01/hour
On-demand
standard
Google Cloud
a2-highgpu-4g
4× A100
40 GB VRAM/GPU · 48 vCPUs · 340 GB RAM · storage unknown
Region: Council Bluffs, Iowa · US · us-central1
  • USD 14.69/hour
    compute time
≈ EUR 12.94/hour
US-CENTRAL1 on-demand
standard
Google Cloud
a2-highgpu-4g
4× A100
40 GB VRAM/GPU · 48 vCPUs · 340 GB RAM · storage unknown
Region: Eemshaven, Netherlands · NL · europe-west4 · EU
  • USD 14.99/hour
    compute time
≈ EUR 13.2/hour
EUROPE-WEST4 on-demand
standard
Lambda
8× NVIDIA A100 SXM · 40 GB/GPU
8× A100
40 GB VRAM/GPU · 124 vCPUs · 1800 GB RAM · 5939.2 GB storage
Region: unknown
  • USD 1.99/hour → USD 15.92/hour effective
    compute time
≈ EUR 14.02/hour
On-demand
standard
Lambda
8× NVIDIA A100 SXM · 80 GB/GPU
8× A100
80 GB VRAM/GPU · 240 vCPUs · 1800 GB RAM · 19968 GB storage
Region: unknown
  • USD 2.79/hour → USD 22.32/hour effective
    compute time
≈ EUR 19.66/hour
On-demand
standard
Google Cloud
a2-highgpu-8g
8× A100
40 GB VRAM/GPU · 96 vCPUs · 680 GB RAM · storage unknown
Region: Council Bluffs, Iowa · US · us-central1
  • USD 29.39/hour
    compute time
≈ EUR 25.88/hour
US-CENTRAL1 on-demand
standard
Google Cloud
a2-highgpu-8g
8× A100
40 GB VRAM/GPU · 96 vCPUs · 680 GB RAM · storage unknown
Region: Eemshaven, Netherlands · NL · europe-west4 · EU
  • USD 29.98/hour
    compute time
≈ EUR 26.41/hour
EUROPE-WEST4 on-demand
standard
Google Cloud
a2-megagpu-16g
16× A100
40 GB VRAM/GPU · 96 vCPUs · 1360 GB RAM · storage unknown
Region: Council Bluffs, Iowa · US · us-central1
  • USD 55.74/hour
    compute time
≈ EUR 49.09/hour
US-CENTRAL1 on-demand
standard
Google Cloud
a2-megagpu-16g
16× A100
40 GB VRAM/GPU · 96 vCPUs · 1360 GB RAM · storage unknown
Region: Eemshaven, Netherlands · NL · europe-west4 · EU
  • USD 56.63/hour
    compute time
≈ EUR 49.87/hour
EUROPE-WEST4 on-demand
standard

Available Providers