{"schema_version":"1.0","generated_at":"2026-08-25T19:04:11.801Z","as_of_date":"2026-08-25","manual_url":"https://cloudgpuprices.com/agents#operational-profiles","interpretation":"Hand-reviewed product behavior. Setup complexity follows the published rubric and is an editorial aid, not a measured benchmark. Exact current billing increments and prices remain in the offerings catalog.","setup_complexity_rubric":{"low":"The common path uses a provider-native SDK or source deployment and does not require separate cloud IAM, quota, and registry setup.","moderate":"The common path requires a provider-specific handler or container plus explicit endpoint, registry, or autoscaling configuration.","high":"The common path spans several cloud resources such as project billing, enabled APIs, IAM, registry or source build, regional selection, and quota."},"query":{"provider":[],"product":[],"setup":[],"review_state":[]},"total_profiles":5,"profiles":[{"id":"beam/serverless","provider":{"slug":"beam","name":"Beam"},"product":{"slug":"serverless","name":"Serverless"},"display_name":"Beam Serverless","scope":{"plans":["Developer","Team","Growth"],"notes":"Applies to Beam functions, endpoints and task queues using serverless containers, not intentionally persistent Pods."},"review":{"verified_at":"2026-07-18","review_after":"2026-10-18","state":"current"},"deployment":{"model":"function_platform","invocation_modes":["synchronous_http","asynchronous_job","task_queue","scheduled_job"],"summary":"Functions, endpoints, task queues and images are defined through Beam's Python SDK and deployed with the Beam CLI. Task queues provide an asynchronous path for long audio and model jobs."},"scaling":{"scale_to_zero":"supported","default_behavior":"default","summary":"Serverless applications scale to zero by default. Configured warm time retains a container after work and is billable."},"billing":{"basis":"container_lifecycle","summary":"GPU, CPU and RAM are additive. Beam bills application startup, execution and configured warm time, while its pricing documentation says infrastructure spin-up and container-image loading are not charged."},"startup":{"summary":"Beam advertises sub-second restore for checkpointed workloads. End-to-end startup still includes application initialization, and enabling checkpoints requires an initial capture and distribution step."},"setup":{"complexity":"low","summary":"The normal workflow uses a Beam account, Python SDK and CLI. Images and infrastructure are declared in Python, with optional custom base images and no required external registry in the common path.","requirements":["Beam account and token","Python SDK and Beam CLI","Function, endpoint or task-queue definition"]},"constraints":["CPU and RAM are material additive charges and should be included when comparing Beam with inclusive machine prices.","Developer-plan GPU concurrency is lower than paid-plan concurrency."],"sources":[{"kind":"overview","url":"https://docs.beam.cloud/v2/getting-started/core-concepts","label":"Beam core concepts"},{"kind":"deployment","url":"https://docs.beam.cloud/v2/reference/py-sdk","label":"Beam Python SDK"},{"kind":"billing","url":"https://docs.beam.cloud/v2/resources/pricing-and-billing","label":"Pricing and billing"},{"kind":"startup","url":"https://www.beam.cloud/pricing","label":"Beam pricing and startup summary"}]},{"id":"google-cloud/cloud-run-gpu","provider":{"slug":"google-cloud","name":"Google Cloud"},"product":{"slug":"cloud-run-gpu","name":"Cloud Run GPU Services"},"display_name":"Google Cloud Run GPU","scope":{"plans":["Cloud Run services","Cloud Run jobs"],"notes":"Applies to GPU-enabled Cloud Run services and jobs using instance-based billing, not Compute Engine GPU virtual machines."},"review":{"verified_at":"2026-07-18","review_after":"2026-10-18","state":"current"},"deployment":{"model":"cloud_container_service","invocation_modes":["synchronous_http","asynchronous_job","scheduled_job"],"summary":"Arbitrary containers can run as request-driven Cloud Run services or finite Cloud Run jobs. Jobs are the direct fit for batch processing; services provide managed HTTP routing and concurrency."},"scaling":{"scale_to_zero":"supported","default_behavior":"execution_scoped","summary":"Services scale to zero when minimum instances is zero; jobs allocate instances only for executions. Setting positive minimum instances keeps fully billed GPU capacity running."},"billing":{"basis":"instance_lifecycle","summary":"GPU workloads require instance-based billing. GPU, CPU and memory are additive and charged for the complete instance lifecycle, including startup and idle time before termination, with the exact minimum and increment shown per pricing variant."},"startup":{"summary":"Google documents GPU-equipped Cloud Run instances starting in approximately five seconds before application initialization. Container imports, model loading and storage access add to the observed cold start."},"setup":{"complexity":"high","summary":"The managed runtime is straightforward after provisioning, but initial setup spans a billed Google Cloud project, enabled APIs, IAM/service identity, an image build or Artifact Registry, regional GPU selection and quota.","requirements":["Google Cloud project with billing","Cloud Run API and required IAM roles","Container image or source build","Regional GPU quota"]},"constraints":["Cloud Run currently permits one supported GPU per instance and enforces minimum CPU and memory sizes for each GPU type.","GPU availability and quotas are regional; zonal redundancy changes the GPU rate for services."],"sources":[{"kind":"overview","url":"https://docs.cloud.google.com/run/docs/configuring/services/gpu","label":"GPU support for services"},{"kind":"deployment","url":"https://docs.cloud.google.com/run/docs/configuring/jobs/gpu","label":"GPU support for jobs"},{"kind":"billing","url":"https://cloud.google.com/run/pricing","label":"Cloud Run pricing"},{"kind":"billing","url":"https://docs.cloud.google.com/run/docs/configuring/billing-settings","label":"Cloud Run billing settings"}]},{"id":"modal/serverless","provider":{"slug":"modal","name":"Modal"},"product":{"slug":"serverless","name":"Serverless Functions"},"display_name":"Modal Serverless Functions","scope":{"plans":["Starter","Team","Enterprise"],"notes":"Applies to Modal Functions and their autoscaling container pools, not continuously provisioned Sandboxes or Notebooks."},"review":{"verified_at":"2026-07-18","review_after":"2026-10-18","state":"current"},"deployment":{"model":"function_platform","invocation_modes":["synchronous_http","asynchronous_job","scheduled_job"],"summary":"Applications, functions, images, secrets and GPU requirements are defined in Python and run with the Modal SDK or deployed with the Modal CLI."},"scaling":{"scale_to_zero":"supported","default_behavior":"default","summary":"Function container pools scale to zero by default. A scaledown window can retain idle containers, while minimum containers deliberately prevents scale-to-zero."},"billing":{"basis":"function_execution","summary":"GPU, CPU and memory are additive usage components. Retaining an idle container through a scaledown window can continue consuming billable resources; exact current meters and increments are shown with each pricing variant."},"startup":{"summary":"Modal documents container boot around one second, but complete cold starts range from seconds to minutes when application imports, model downloads or initialization are included. Images, volumes and memory snapshots can reduce application warm-up time."},"setup":{"complexity":"low","summary":"The common path needs a Modal account, Python package and CLI authentication; infrastructure configuration and image construction remain in Python and a separate container registry is not required.","requirements":["Modal account","Python SDK and Modal CLI","Application code expressed as a Modal App"]},"constraints":["Selecting or constraining a compute region can apply a regional price multiplier.","Total cost depends on requested CPU and memory as well as the GPU rate."],"sources":[{"kind":"overview","url":"https://modal.com/docs/guide","label":"Modal introduction"},{"kind":"deployment","url":"https://modal.com/docs/guide/apps","label":"Apps and deployments"},{"kind":"scaling","url":"https://modal.com/docs/guide/scale","label":"Autoscaling"},{"kind":"startup","url":"https://modal.com/docs/guide/cold-start","label":"Cold-start performance"},{"kind":"billing","url":"https://modal.com/pricing","label":"Modal pricing"}]},{"id":"ovhcloud/ai-deploy","provider":{"slug":"ovhcloud","name":"OVHcloud"},"product":{"slug":"ai-deploy","name":"AI Deploy"},"display_name":"OVHcloud AI Deploy","scope":{"plans":["Public Cloud AI Deploy"],"notes":"Applies to AI Deploy application replicas, not AI training jobs, notebooks or Public Cloud GPU virtual machines."},"review":{"verified_at":"2026-07-18","review_after":"2026-10-18","state":"current"},"deployment":{"model":"managed_container","invocation_modes":["synchronous_http"],"summary":"A Docker image is deployed as a managed HTTP application with one or more GPU-backed replicas through the OVHcloud control panel or ovhai CLI."},"scaling":{"scale_to_zero":"conditional","default_behavior":"requires_configuration","summary":"Scale-to-zero requires autoscaling with minimum replicas set to zero. The scale-down and scale-to-zero stabilization windows determine how long the last replica remains running."},"billing":{"basis":"instance_lifecycle","summary":"Compute is charged for each replica's lifetime with per-minute granularity, rather than per request. Registry and persistent storage costs are separate."},"startup":{"summary":"OVHcloud documents a 30-second-to-several-minute cold-start delay from zero depending on image and volume weight, with an additional risk that a popular GPU flavor is temporarily unavailable."},"setup":{"complexity":"moderate","summary":"Deployment requires an OVHcloud Public Cloud project, a compatible container image in a reachable registry, selected AI Deploy compute resources and explicit autoscaling configuration for scale-to-zero.","requirements":["OVHcloud Public Cloud project","Container image and registry","AI Deploy application configuration","Autoscaling minimum zero for scale-to-zero"]},"constraints":["The product is an HTTP application platform rather than a native queued batch-job abstraction.","A request arriving at zero replicas waits through the complete cold start, and capacity is not reserved."],"sources":[{"kind":"deployment","url":"https://help.ovhcloud.com/csm/en-au-public-cloud-ai-deploy-getting-started?id=kb_article_view&sysparm_article=KB0047969","label":"AI Deploy getting started"},{"kind":"scaling","url":"https://help.ovhcloud.com/csm/en-ie-public-cloud-ai-deploy-apps-deployments?id=kb_article_view&sysparm_article=KB0048008","label":"AI Deploy scaling strategies"},{"kind":"billing","url":"https://help.ovhcloud.com/csm/it-public-cloud-ai-deploy-billing?id=kb_article_view&sysparm_article=KB0057061","label":"AI Deploy billing and lifecycle"},{"kind":"billing","url":"https://www.ovhcloud.com/en/public-cloud/prices/","label":"OVHcloud Public Cloud pricing"}]},{"id":"runpod/serverless","provider":{"slug":"runpod","name":"Runpod"},"product":{"slug":"serverless","name":"Serverless"},"display_name":"Runpod Serverless","scope":{"plans":["Flex workers"],"notes":"Describes scale-to-zero Flex workers. Active workers remain running and have different cost and latency behavior."},"review":{"verified_at":"2026-07-18","review_after":"2026-10-18","state":"current"},"deployment":{"model":"queued_worker","invocation_modes":["synchronous_http","asynchronous_job","task_queue"],"summary":"A provider-specific handler runs inside a custom container behind a managed queue endpoint. Images can be supplied from a container registry or deployed from a GitHub repository."},"scaling":{"scale_to_zero":"supported","default_behavior":"default","summary":"Flex workers scale dynamically to zero when idle. Active workers are a separate always-running option for workloads that must avoid cold starts."},"billing":{"basis":"container_lifecycle","summary":"Flex workers are billed from worker startup until complete shutdown, including initialization and the configured idle timeout, rounded to the provider's billing increment. Storage is charged separately."},"startup":{"summary":"Cold-start time depends on image download, worker initialization and model loading. Cached models and FlashBoot can reduce it, but the general worker path has no single guaranteed end-to-end cold-start time."},"setup":{"complexity":"moderate","summary":"The standard workflow requires a Runpod handler, local container testing, an image registry or GitHub deployment, and endpoint/GPU configuration in the Runpod console or API.","requirements":["Runpod account","Runpod handler integration","Docker image or GitHub repository","Serverless endpoint configuration"]},"constraints":["Several lower-cost GPU classes group multiple GPU models, so exact hardware and performance may vary within the selected class.","Worker concurrency limits depend on account balance unless increased by support."],"sources":[{"kind":"overview","url":"https://docs.runpod.io/serverless/overview","label":"Serverless overview"},{"kind":"deployment","url":"https://docs.runpod.io/serverless/quickstart","label":"Serverless quickstart"},{"kind":"scaling","url":"https://docs.runpod.io/serverless/workers/overview","label":"Worker types and lifecycle"},{"kind":"billing","url":"https://docs.runpod.io/serverless/pricing","label":"Serverless billing"}]}]}