Modal Serverless GPUs in Production: Cold Starts, Concurrency, and Costs
Why We Brought This Tool Into Our Lab
We brought Modal into our evaluation queue for one reason: most GPU platforms make bursty inference economically awkward.
A 15-second cold path on 1,000 daily A100 jobs adds $312.30, or 25%, to 30-day GPU compute before CPU and memory.
A conventional GPU deployment asks us to choose between two bad outcomes. We can keep enough capacity online for peak traffic and pay for idle accelerators, or provision closer to average demand and watch queues and tail latency explode during a spike. Kubernetes does not remove that tradeoff. It gives us more machinery with which to operate it.
Modal approaches the problem at a higher level. We define a Python function or class, select a GPU, and let the platform create and retire containers around demand. The attractive part is not the decorator syntax. It is the prospect of paying for short-lived GPU allocation rather than operating an always-on pool.
We identified four mechanisms that support this approach:
- A shared buffer of healthy GPU machines removes cloud instance acquisition from the immediate request path.
- A content-addressed, lazy-loading filesystem avoids downloading an entire container image before starting the process.
- CPU checkpoint and restore can skip repeated host-side initialization.
- CUDA checkpoint and restore can preserve device-side state, including initialized CUDA contexts.
The filesystem design is particularly relevant for oversized ML images. In our filesystem review, we identified a metadata-only startup path involving a few megabytes and a design reference of 100 milliseconds or less, with accessed files loaded afterward on demand; we did not measure that timing locally. That does not make model weights free to load, but it removes a large amount of irrelevant image data from the critical path.
In our architecture review, we used the improvement from roughly 2,000 seconds to about 50 seconds as a representative reference for complex inference-server startup, not as a timing measurement from our environment. We did not treat it as a universal cold-start result or a guarantee for every model, image, region, or GPU. Our takeaway was its order of magnitude: optimized replicas may arrive in seconds or tens of seconds instead of the tens of minutes associated with naïve instance provisioning.
That distinction matters. A 50-second cold start is excellent for a batch worker and unacceptable for an interactive request with a two-second latency objective. “Serverless” does not mean “startup latency disappears.”
We could not run a live GPU test because a live Modal deployment and actual billing records were unavailable in our review environment. GPU execution also requires a valid payment method. We reviewed the documented interface, prepared an illustrative probe, and worked through pricing estimates, but we did not execute the probe, verify a version-pinned SDK contract, or measure cloud startup and concurrency behavior.
That limitation narrows our verdict rather than invalidating the review. Modal's architecture and pricing are still concrete enough to determine where it fits, while the included, untested probe provides a starting point for account-specific latency testing.
Hands-On Walkthrough: Setup, Execution & Output
We prepared the following setup instructions using a Modal SDK major-version range. This is not an exact version pin, and we did not execute or validate the setup in our review environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "modal>=1,<2"
modal setup
modal profile current
modal setup opens the authentication flow and stores the workspace credentials locally. A GPU function will still fail if the workspace lacks a valid payment method or sufficient GPU concurrency.
We then prepared this probe as modal_gpu_probe.py:
import json
import socket
import time
import modal
app = modal.App("effloow-modal-gpu-probe")
image = (
modal.Image.debian_slim(python_version="3.12")
.pip_install("torch==2.7.1")
)
@app.cls(
image=image,
gpu="A100-80GB",
timeout=300,
max_containers=8,
)
@modal.concurrent(max_inputs=1)
class GPUProbe:
@modal.enter()
def initialize(self):
import torch
self.torch = torch
self.ready_at = time.time()
# Force CUDA context creation during container initialization.
torch.cuda.init()
torch.cuda.synchronize()
@modal.method()
def run(self, matrix_size: int = 2048):
torch = self.torch
started = time.perf_counter()
left = torch.randn(
matrix_size,
matrix_size,
device="cuda",
dtype=torch.float16,
)
right = torch.randn(
matrix_size,
matrix_size,
device="cuda",
dtype=torch.float16,
)
output = left @ right
torch.cuda.synchronize()
kernel_ms = (time.perf_counter() - started) * 1000
return {
"container": socket.gethostname(),
"gpu": torch.cuda.get_device_name(0),
"matrix_size": matrix_size,
"kernel_ms": round(kernel_ms, 2),
"output_shape": list(output.shape),
"seconds_since_ready": round(time.time() - self.ready_at, 3),
}
@app.local_entrypoint()
def main(burst: int = 8):
probe = GPUProbe()
cold_started = time.perf_counter()
first = probe.run.remote(2048)
cold_e2e_seconds = time.perf_counter() - cold_started
burst_started = time.perf_counter()
results = list(
probe.run.map(
[2048] * burst,
order_outputs=False,
)
)
burst_e2e_seconds = time.perf_counter() - burst_started
summary = {
"cold_e2e_seconds": round(cold_e2e_seconds, 3),
"burst_requests": burst,
"burst_e2e_seconds": round(burst_e2e_seconds, 3),
"containers_observed": len(
{result["container"] for result in results}
),
"first_result": first,
}
print(json.dumps(summary, indent=2))
The command is straightforward:
modal run modal_gpu_probe.py --burst 8
For repeatable cold-start testing, we would allow every container to scale down between runs, execute at least 30 trials, and retain p50, p95, and p99 client-observed latency. For warm behavior, we would invoke the same deployment repeatedly without crossing the scale-down window. We would also record container identities because eight simultaneous requests do not prove that eight replicas were created.
The following is simulated output showing the intended JSON shape, not a result from a live Modal run:
{
"cold_e2e_seconds": 8.417,
"burst_requests": 8,
"burst_e2e_seconds": 12.806,
"containers_observed": 5,
"first_result": {
"container": "modal-ctnr-7f31",
"gpu": "NVIDIA A100-SXM4-80GB",
"matrix_size": 2048,
"kernel_ms": 3.84,
"output_shape": [2048, 2048],
"seconds_since_ready": 0.391
}
}
Those numbers exist only to demonstrate parsing and dashboard ingestion. They must not be used as a Modal benchmark. The client-observed cold duration includes allocation, image startup, imports, CUDA initialization, scheduling, network transit, and execution. The in-container kernel duration includes almost none of that.
For production inference, we would replace the matrix operation with model initialization in @modal.enter() and inference in @modal.method(). We would store weights on a Modal Volume rather than downloading them from a public model registry on every boot. For startup planning, we used a Volume read range of 1–2 GB/s, not a locally measured result. At that rate, each gigabyte of weights adds roughly half a second to a second of loading time before framework initialization and compilation.
We would then consider memory snapshots only after establishing an unsnapshotted baseline. In our snapshot review, we used roughly 10x cold-start reductions as a reference for suitable workloads, not as a result measured in our environment. We would still check snapshot compatibility for the specific application. CUDA graph capture, JIT kernels, and framework compilation can still turn startup into a tens-of-seconds or even minutes-long path if they are repeated on every replica.
Account Requirements, Operational Limits, and Pricing Risks
Our practical limitation was the absence of a live Modal deployment and actual billing records. Python ergonomics do not bypass authentication, payment, or workspace limits. We would use the following untested authentication preflight in CI, while checking payment status and GPU concurrency separately:
set -euo pipefail
modal profile current >/dev/null || {
echo "Modal authentication is missing" >&2
exit 1
}
python -c "import modal; print(modal.__version__)"
The second problem was definitional. “Cold start” can refer to container scheduling, Python process startup, CUDA initialization, model loading, server readiness, or the complete client request. Mixing these measurements creates impressive but useless charts. Our harness keeps client-observed end-to-end time separate from in-container work. A production harness should additionally timestamp model-ready state and first-token latency.
The third gotcha is that a warm container is not the same as free idle capacity. Modal's base proposition is scale-to-zero with no GPU charge after resources are released. Keeping application-level buffer containers available to absorb spikes increases cost. We would not approve a budget based on “no idle charges” without checking the actual invoice behavior for the selected warm-pool and scale-down settings.
Cold starts are billable work. If loading a model takes 15 seconds, that time belongs in the GPU-second estimate even though it produces no inference output.
We also found several hardware-selection traps:
gpu="A100"can be upgraded from a 40 GB to an 80 GB A100 without changing the requested price. We would useA100-40GBorA100-80GBwhen hardware identity matters.gpu="H100"can be upgraded to an H200. We would useH100!for a controlled H100 benchmark.- Requests for more than two GPUs in one container generally face longer allocation waits.
- B300 requires CUDA 13.1 or later, so an image tested on Hopper is not automatically Blackwell Ultra compatible.
- A faster GPU can cost more without improving a memory-bound, batch-one decode workload enough to justify it.
Concurrency also needs deliberate tuning. @modal.concurrent exposes max_inputs and target_inputs, but raising concurrency does not create free throughput. Too many simultaneous model inputs can exhaust KV-cache memory, increase queue time inside the engine, and make p99 latency worse while the GPU looks busy. For latency-sensitive LLM serving, we would tune engine-level continuous batching together with Modal's container target rather than treating them as independent controls.
Finally, the pricing multipliers can dominate GPU choice. In our September 2026 rate audit, a broad pinned region introduced a 1.5x multiplier, a narrower region could reach 1.75x, and non-preemptible execution carried a 3x multiplier. A narrow region combined with non-preemptible execution can therefore reach 5.25x the headline rate. These are published mechanics, but we would still re-check the current Modal pricing page before signing a budget.
Network egress was not sufficiently itemized in the pricing material we reviewed. We treat it as an unknown, not as zero. That matters for video generation, image pipelines, and long streamed responses.
Scale, Latency & Cost vs. Alternatives
The audited September 2026 base rates were $0.000694 per second for an A100 80 GB and $0.001097 per second for an H100 SXM5. Those convert to approximately $2.50 and $3.95 per GPU-hour, before CPU, memory, region, preemption, storage, tax, and plan-level charges.
Our cost model separates compute from storage and plan fees. In the formula below, we apply region and execution multipliers only to the GPU, CPU, and memory compute subtotal, then add storage and plan charges. The final line refers only to that compute subtotal, not the entire bill:
Monthly cost =
GPU seconds × GPU rate
+ CPU core-seconds × CPU rate
+ memory GiB-seconds × memory rate
+ storage and plan charges
Then apply any applicable region and execution multipliers.
Consider 1,000 A100 jobs per day, each doing 60 seconds of useful work and incurring a 15-second cold path:
GPU seconds per day = 1,000 × (60 + 15) = 75,000
Daily A100 cost = 75,000 × $0.000694 = $52.05
30-day A100 cost = $1,561.50
Without the cold path, the same useful work costs $1,249.20. Repeated startup adds $312.30, or 25%, before CPU and memory. Batching ten jobs into each container invocation would sharply reduce that penalty if the workload tolerates queueing.
The same 75,000 daily GPU-seconds at the H100 base rate would cost about $82.28 per day, or $2,468.25 over 30 days. H100 is cheaper only if its measured throughput or latency improvement offsets the higher rate. GPU generation is not a cost optimization by itself.
| Option | Billing and scaling model | Cold-path exposure | Operational burden | Best fit |
|---|---|---|---|---|
| Modal | Per-second serverless allocation; scale to zero | Seconds to tens of seconds for optimized replicas; heavy initialization can take minutes | Low to moderate | Bursty inference, parallel batch work, development |
| Reserved cloud GPU | Fixed capacity for a contract or allocation period | Usually hidden by always-on capacity | Moderate | Stable, continuously utilized demand |
| Self-managed Kubernetes GPU pool | Pay for nodes while provisioned | Low when nodes and pods remain warm | High | Teams needing deep scheduler and network control |
| Managed model API | Per token, image, or output | Abstracted from the customer | Low | Commodity models where infrastructure ownership adds little value |
| Hybrid baseline plus burst | Reserved baseline with serverless overflow | Applies mainly to burst replicas | High because two systems remain | Large workloads with predictable floor and volatile peaks |
The practical break-even equation is more useful than a generic “serverless is cheaper” statement:
Serverless cost = serverless rate × average demand × time
Reserved cost = reserved rate × peak provisioned demand × time
Serverless wins when:
peak demand / average demand > serverless rate / reserved rate
If a reservation is five times cheaper per GPU-hour, serverless needs a peak-to-average ratio above roughly 5:1 to win on raw accelerator cost. In our break-even analysis, we considered inference, training, and agentic workload ratios around 5:1 to 10:1, and excess burst-load scenarios around 10:1 to 100:1; we did not measure those ratios from a deployment of our own. Those are plausible workload shapes, but we would calculate the ratio from our own production trace before making a procurement decision.
At high steady utilization, the advantage reverses. If a dedicated H100 costs 75% of Modal's effective hourly rate, the dedicated unit becomes cheaper at approximately 75% sustained utilization, ignoring staffing and platform costs. Region and non-preemptible multipliers lower that threshold further.
Modal's strongest position is therefore not “cheapest GPU.” It is avoiding payment for idle capacity during periods of low demand. It also removes cluster upgrades, node health checks, autoscaler integration, and much of the capacity brokerage. Those engineering savings are real, but we would account for them separately from raw compute costs.
Teams evaluating this alongside a broader stack can review our AI infrastructure tools or talk to us about a workload-specific cost model. We would use at least four weeks of request traces, including burst duration and geographic constraints, rather than extrapolating from average tokens per month.
Our Final Verdict: When to Deploy, When to Skip
Modal is a credible serverless GPU platform for workloads whose demand genuinely falls toward zero or arrives in large, parallel bursts. Its technical value comes from moving machine acquisition out of the request path, lazily starting container filesystems, and offering CPU and GPU snapshot mechanisms. The Python interface is the easy part; the shared capacity and startup engineering are the product.
We would deploy it when:
- Traffic has a measured peak-to-average ratio above the effective serverless-to-reserved price ratio.
- Jobs are asynchronous, retryable, or tolerant of cold starts measured in seconds.
- Batch work can be grouped to amortize model loading.
- The team does not want to operate GPU nodes, drivers, health checks, and cluster autoscaling.
- The model fits on one or two GPUs and can use a broad capacity pool.
- Scale-to-zero savings outweigh the need for a permanent warm replica.
- The application can use snapshots or cached Volumes to reduce initialization.
- We can remain preemptible and avoid narrow regional pinning.
We would hold off or avoid it when:
- Every request has a hard sub-second response objective and traffic is too sparse to justify a warm pool.
- GPU utilization remains high and predictable for most of the month.
- Compliance requires a narrow region and guaranteed non-preemptible execution without room for the resulting multiplier.
- The deployment needs more than two tightly coupled GPUs per replica and allocation delay is critical.
- The workload depends on unsupported kernels, a precise GPU SKU, custom networking, or low-level host access.
- Finance requires a fully verified egress schedule before deployment.
- The organization already operates a highly utilized reserved GPU fleet with competent on-call coverage.
Our final assessment is positive but conditional: Modal solves allocation waste better than it solves latency. For batch processing and volatile inference, that is often exactly the right trade. For a continuously busy model server, per-second billing can become an expensive way to rent a GPU that should have been reserved.
Before production approval, we would run the included probe in the target region, replace the matrix operation with the real model, collect at least 30 true cold starts, and reconcile one week of measured GPU-seconds against the invoice. We would then compare that result with a reserved baseline using the same throughput and SLO—not an assumed tokens-per-second figure.
For architecture help beyond the initial proof of concept, see our AI infrastructure services. For the underlying platform design and economic model, we also recommend reading Modal's detailed explanations of serverless GPU startup and serverless GPU pricing.
Get the next one
in your inbox.
One short weekly dispatch with new guides, tools, and what we tested. No spam, unsubscribe anytime.
Get weekly AI tool reviews & automation tips
Join our newsletter. No spam, unsubscribe anytime.
More in Articles
We favor Docker Model Runner for Docker-centered team workflows and Ollama for simpler standalone setup, without declaring a cold-start or memory winner.
We deployed Langfuse and Arize Phoenix locally, pushed synthetic LLM traces through their OpenTelemetry interfaces, exercised evaluation workflows, and measured the operational cost hidden behind self-hosting.
We deployed Windmill with Docker Compose, built a mixed-language data and AI workflow, tested worker isolation, and mapped the operational trade-offs against Airflow.
We deployed FastMCP behind Docker and nginx, wired it to OAuth token introspection, and tested Streamable HTTP session handling, proxy headers, callback URLs, and DNS rebinding defenses.
Tools you can use
Free calculator: compare the monthly cost of an LLM API against self-hosting on your own GPU, and find the token volume where self-hosting starts to win.
Free calculator: estimate the GPU VRAM needed to run or fine-tune any LLM. Choose model size, precision, and mode (inference, LoRA, QLoRA, full fine-tune).