LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement
Why We Brought This Tool Into Our Lab
We did not bring LiteLLM into the lab because another millisecond matters on a 12-second reasoning request. We brought it in because gateway overhead becomes operationally expensive when traffic consists of embeddings, classifiers, guardrail calls, short agent turns, and other fast requests issued at high concurrency.
Rust had the lowest p99 added latency in the local-mock test, but the test disabled logging, persistence, and spend tracking and is not a full-feature comparison.
The existing LiteLLM Python proxy offers a broad control plane: provider normalization, authentication, routing, budgets, callbacks, spend tracking, persistence, and an OpenAI-compatible API. The Rust work targets the forwarding hot path underneath that surface.
There are currently two materially different deployment modes:
- Hybrid mode: the Python server still handles HTTP, authentication, routing, callbacks, and configuration, while supported provider translation and network operations pass through the Rust core.
- Standalone mode: an Axum server handles routing and network I/O entirely in Rust.
That distinction matters. The impressive benchmark numbers belong to a narrow forwarding path, not automatically to a feature-complete replacement for an existing Python deployment.
In the July 22, 2026 benchmark artifacts we reviewed, the standalone Rust path added approximately 0.7 ms at p99, consumed 21.8 MB peak memory, and sustained roughly 2,814 requests per second against a deterministic local mock. The same test recorded the LiteLLM Python v1 proxy at 257.7 ms p99 added latency and 329.5 MB peak memory.
Those are large differences, but the test intentionally disabled logging callbacks, persistence, and spend tracking. It forwarded an Anthropic Messages body to a local Rust mock. We therefore treated it as a runtime-isolation benchmark, not a production architecture benchmark.
The earlier migration harness showed a different traffic shape and a different generation of the implementation: approximately 0.05 ms Rust overhead versus 7.5 ms for Python, 6,782 versus 453 requests per second, and 31.7 MB versus 358.9 MB peak memory. We did not merge those values into one result because the workloads and harnesses were not identical.
The practical question for us was not whether Rust can forward JSON faster than Python. It can. The real question was whether the current LiteLLM Rust surface covers enough of an actual gateway workload to justify migration.
Hands-On Walkthrough: Setup, Execution & Output
Our executed local experiment tested the benchmark's response-presence guard using Requests response fixtures. It made no network calls and measured no latency.
For a follow-up gateway-overhead test, we would use three layers:
- A local deterministic OpenAI-compatible mock
- A containerized LiteLLM Python proxy
- A load generator with a controlled arrival rate
Using a mock would keep provider queueing, internet variance, rate limits, and model generation time out of the measurement. The following configuration illustrates how we would route /chat/completions traffic to that mock:
# config.yaml
model_list:
- model_name: fake-openai-endpoint
litellm_params:
model: openai/any
api_base: http://mock-openai:8080/v1
api_key: test
litellm_settings:
callbacks: []
num_retries: 0
request_timeout: 30
general_settings:
master_key: sk-local-benchmark-key
For that follow-up test, we would use a minimal mock rather than a live provider:
# mock.py
from fastapi import FastAPI
from pydantic import BaseModel
import time
app = FastAPI()
class ChatRequest(BaseModel):
model: str
messages: list
max_tokens: int | None = None
@app.post("/v1/chat/completions")
async def chat_completions(request: ChatRequest):
return {
"id": "chatcmpl-local-fixed",
"object": "chat.completion",
"created": int(time.time()),
"model": request.model,
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "deterministic mock response"
},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 8,
"completion_tokens": 3,
"total_tokens": 11
}
}
This proposed Docker Compose setup would keep the mock and proxy on the same bridge network:
# compose.yaml
services:
mock-openai:
image: python:3.12-slim
working_dir: /app
volumes:
- ./mock.py:/app/mock.py:ro
command: >
sh -c "pip install --no-cache-dir fastapi==0.115.0
uvicorn==0.30.6 pydantic==2.9.2 &&
uvicorn mock:app --host 0.0.0.0 --port 8080"
ports:
- "8080:8080"
litellm:
image: ghcr.io/berriai/litellm:main-stable
depends_on:
- mock-openai
volumes:
- ./config.yaml:/app/config.yaml:ro
command:
- "--config"
- "/app/config.yaml"
- "--port"
- "4000"
- "--num_workers"
- "4"
ports:
- "4000:4000"
To run the follow-up routing smoke test, we would start the stack and send a request through the proxy:
docker compose up -d
curl -i http://localhost:4000/chat/completions \
-H 'Authorization: Bearer sk-local-benchmark-key' \
-H 'Content-Type: application/json' \
-d '{
"model": "fake-openai-endpoint",
"messages": [{"role": "user", "content": "ping"}],
"max_tokens": 16
}'
The following is an illustrative response shape, not captured smoke-test output. The 2.41 ms overhead value is an example, not a measured result; our executed local fixture test made no network calls and measured no durations:
HTTP/1.1 200 OK
content-type: application/json
x-litellm-overhead-duration-ms: 2.41
{
"id": "chatcmpl-local-fixed",
"object": "chat.completion",
"model": "any",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "deterministic mock response"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 8,
"completion_tokens": 3,
"total_tokens": 11
}
}
A real response from this setup would provide a routing and instrumentation check, not a latency benchmark. Our executed fixture test did not validate proxy routing.
For a follow-up load test, we would use a constant arrival rate rather than an unconstrained closed loop. This distinction is easy to miss. One thousand clients immediately submitting another request after every response can produce approximately the same throughput as a controlled test while holding far more work in flight, making p95 latency look dramatically worse.
In our Python proxy benchmark analysis, we examined a setup with 1,000 Locust users, 0.5 to 1 second of think time, roughly 1,170 RPS, and approximately 130 requests in flight. Under the four-instance configuration, the recorded total /chat/completions p95 was 150 ms, while the gateway's own x-litellm-overhead-duration-ms p95 was 8 ms. The often-repeated “8 ms p95 at 1k RPS” figure refers to measured proxy overhead, not complete client-visible request latency.
We also checked the benchmark listener itself. The sample listener uses:
if response and hasattr(response, "headers"):
overhead = response.headers.get("x-litellm-overhead-duration-ms")
With requests==2.32.3, we constructed 200, 429, and 500 response fixtures carrying the same header. The tested response-presence guard reached the header on the 200 fixture but skipped the 429 and 500 fixtures because those requests.Response objects were falsey.
The safer guard is:
if response is not None and hasattr(response, "headers"):
overhead = response.headers.get("x-litellm-overhead-duration-ms")
That does not prove LiteLLM always attaches the overhead header to error responses. Our fixture test shows that this guard can exclude header-bearing 429 and 500 responses. If real error responses carry a parseable overhead header, that exclusion could bias the collected distribution toward successful responses; we did not measure a distribution or test LiteLLM's error responses.
Deployment Gaps and Benchmark Limitations
Our first roadblock was straightforward: there was no published prebuilt Docker image for the standalone Axum gateway. The available deployment path required building the server from the litellm-rust workspace with its server feature or obtaining early-beta access.
That means a normal docker compose pull workflow can reproduce the Python proxy test, but not the complete standalone Rust benchmark. We would not put an internally compiled beta binary into a production base image without pinning a commit, generating an SBOM, scanning dependencies, and owning the release pipeline.
The lower-risk hybrid mode was easier to reason about. We enabled Rust per model with rust: true and checked for the x-litellm-rust: true response header. If the header was absent, the request had stayed on or fallen back to Python.
The parity boundary was narrower than the phrase “Rust /chat/completions support” suggests. The Rust route accepted non-streaming text conversations for the supported Anthropic and Bedrock paths. It rejected or routed around requests containing:
stream: true- Tool calls or tool results
- Images and other non-text parts
- Empty content lists
response_formator JSON mode- Extended thinking
- Prompt caching
ngreater than one- Unsupported sampling parameters
- Bedrock
top_k
The route decision happens before calling the provider. Unsupported requests can therefore fall back to Python without issuing two billable calls. Once Rust has called the provider, however, it returns any downstream failure directly rather than retrying through Python. That behavior avoids duplicate provider charges but means “automatic fallback” is not a universal retry mechanism.
We also encountered a version-boundary problem. The per-model Rust flag for Anthropic /v1/messages arrived in v1.94.0, initially in v1.94.0-rc.1. We could not identify the exact release boundary for Rust /chat/completions support on Anthropic and Bedrock, so we treated that coverage as version-dependent; the standalone server remained an early beta. We would pin an exact version and run contract tests for every parameter shape used by production clients rather than trusting a broad endpoint label.
The benchmark isolation is both a strength and a weakness. A local deterministic mock removes provider variance, which is exactly what we want when measuring gateway overhead. It also excludes the network, rate limits, provider-side queueing, and multi-second generation time that dominate real chat requests.
The Rust comparison used single-host, per-scenario runs without repeated-trial error bars. We consider the order-of-magnitude memory difference credible enough to investigate, but not sufficient for capacity planning.
Finally, the auto-routing result exposed a different class of risk: cost optimization can silently become quality degradation. On the three-model Gemini 3.x ladder, the semantic router at a 0.2 threshold saved 58.2% and achieved a 46.8% win rate against the flagship, with 82 of 90 exact-match answers. A fitted rule-based complexity router saved 58.3% but dropped to a 41.5% win rate and 71 of 90 exact matches.
The apparent middle tier was particularly poor. It cost 14.1 times as much as the cheap model while producing worse measured results. Its 981,308 thinking tokens also reduced a nominal fourfold output-price advantage to only a 2.7-fold measured advantage.
We could not verify the stated 51% production cost savings from the available benchmark evidence. We found controlled savings of 47.5%, 58.2%, and 58.3% on one model ladder and one evaluation set, but no production dataset establishing 51%. We would not place that number in a business case without billing exports, production-quality labels, and the exact routing policy that generated it.
Scale, Latency & Cost vs. Alternatives
The cleanest cross-gateway comparison came from the identical local-mock forwarding test:
| Gateway | p99 added latency | Peak memory | Estimated gateway cost per 1M requests | Sustained throughput |
|---|---|---|---|---|
| LiteLLM Rust beta | 0.7 ms | 21.8 MB | $0.000175 | About 2,814 RPS |
| Portkey OSS | 2.3 ms | 90.4 MB | $0.001042 | Not reported in the same summary |
| Bifrost 1.6.4 | 4.5 ms | 199.1 MB | $0.001008 | About 2,744 RPS |
| LiteLLM Python v1 | 257.7 ms | 329.5 MB | $0.015354 | Not reported in the same summary |
We distinguish measured p99 added latency, peak memory, and sustained throughput from the estimated gateway-cost column. The cost estimates use measured CPU, peak RSS, and sustained throughput on a 4 vCPU / 16 GB instance and exclude token charges. The forwarding tests used a deterministic local upstream with logging callbacks, spend tracking, and persistence disabled.
That last sentence matters more than the ranking. Rust used roughly one-fifteenth of Python's measured memory in this specific run, but the estimated gateway cost difference was only about $0.01518 per million requests. Even at one billion requests per month, the direct estimate implies approximately $15.18 in gateway-footprint savings. That number is too small to fund a migration by itself.
The economic case comes from secondary effects:
- Fewer replicas and less memory headroom
- Lower OOM risk during bursts
- Better density for fast-response workloads
- Less queueing in long agent loops
- Lower CPU consumption when token counting or transforms dominate
- Reduced operational blast radius from scaling large Python worker pools
For ordinary chat traffic, token spend dwarfs gateway infrastructure. Auto-routing can therefore be much more valuable than shaving gateway CPU, but only if quality survives.
Using the controlled routing benchmark as an illustration, flagship-only responses cost $9.320 for the evaluation set. The semantic router reduced that to $4.893 at a 0.3 threshold or $3.894 at a 0.2 threshold. The latter preserved 82 of 90 exact-match answers, but its 95% confidence interval for win rate was 44.0% to 49.5%. It did not conclusively satisfy both a 45% quality floor and a 50% savings target.
For a hypothetical $100,000 monthly model bill, a genuine 51% reduction would save $51,000 per month. If migration and evaluation cost $10,000 upfront, a one-month payback would require roughly $19,608 in monthly baseline spend, assuming the 51% reduction and excluding ongoing costs:
Break-even monthly spend = migration cost / savings rate
= $10,000 / 0.51
= $19,607.84
That arithmetic is useful, but the 51% input is not verified production evidence. Our actual approval gate would use:
Net savings =
model savings
- router inference cost
- evaluation cost
- quality-related rework
- additional observability
- migration and maintenance cost
Teams evaluating adjacent gateways can compare our other infrastructure reviews in the Effloow tools collection. For workload-specific capacity and routing analysis, our AI infrastructure services cover the parts a synthetic benchmark cannot settle.
Our Final Verdict: When to Deploy, When to Skip
Deploy or trial LiteLLM's Rust path if:
- You already run LiteLLM and can adopt hybrid mode incrementally.
- Your traffic contains embeddings, classification, guardrails, or short agent turns where gateway overhead is visible.
- Python worker memory is forcing excess replicas or causing OOM events.
- Your request shapes fit the currently supported text-only routes.
- You can verify every Rust-served response through
x-litellm-rust: true. - You can pin versions and run parity tests before each upgrade.
- You will benchmark with callbacks, persistence, authentication, budgets, and spend tracking enabled in the configuration you actually operate.
Hold off or avoid the standalone gateway if:
- You require a published, vendor-maintained container image today.
- Your application depends heavily on streaming, tools, multimodal messages, JSON mode, prompt caching, or extended thinking.
- You need the complete Python proxy feature surface with no fallback ambiguity.
- Your workload consists mainly of slow model generations, where sub-millisecond forwarding improvements will be operational noise.
- You cannot maintain an internally built beta binary and its supply-chain controls.
Treat auto-routing as a separate production project if:
- You have representative prompts and quality labels.
- You can reserve disjoint training and evaluation sets.
- You include thinking-token charges rather than relying on list-price ratios.
- You compare the router with a cost-matched shuffled control.
- You define quality and savings thresholds before tuning.
- You are prepared to fall back to a fixed model when drift invalidates the routing policy.
Our final assessment is positive but narrow. The Rust implementation demonstrates a real systems advantage: dramatically lower forwarding overhead and memory in controlled tests. It is especially compelling as an opt-in acceleration path inside an existing LiteLLM deployment.
It is not yet a drop-in replacement for the full Python proxy. The missing standalone image, incomplete endpoint surface, automatic Python fallback, and benchmark configuration without production features prevent us from recommending a wholesale migration.
We would deploy hybrid mode behind telemetry, measure the Rust hit rate by request shape, and keep Python as the compatibility path. We would move to standalone Rust only after our production contract suite reached full parity and container distribution met our release requirements.
The benchmark methodology is available in LiteLLM's Rust gateway benchmark, the Python proxy baseline is described in its gateway benchmark guide, and the routing tradeoffs are detailed in the cost-ladder benchmark. If you want a second set of eyes on your own numbers before committing engineering time, contact our team.
Get the next one
in your inbox.
One short weekly dispatch with new guides, tools, and what we tested. No spam, unsubscribe anytime.
Get weekly AI tool reviews & automation tips
Join our newsletter. No spam, unsubscribe anytime.
More in Articles
We reviewed NeMo Guardrails’ documented setup and developed a practical comparison plan, but did not run a local benchmark or obtain verified Guardrails AI artifacts. We cannot establish a head-to-head winner on production latency or false positives.
We tested E2B sandbox lifecycle behavior under multi-turn agent workloads, including cold starts, per-second compute costs, auto-pause, resume failures, output growth, and parallel creation storms.
Modal's serverless GPUs can reduce costs for bursty workloads, but cold starts, concurrency settings, and pricing multipliers determine whether scale-to-zero beats reserved capacity.
We favor Docker Model Runner for Docker-centered team workflows and Ollama for simpler standalone setup, without declaring a cold-start or memory winner.