Articles, one at a time.
Every piece cites primary sources and adds something you can act on: a worked example, a tested command, a documented limitation. Drafts that only summarize what already exists stay unpublished.
FastMCP Streamable HTTP in Production: Auth, DNS Rebinding, and Docker Gotchas
We deployed FastMCP behind Docker and nginx, wired it to OAuth token introspection, and tested Streamable HTTP session handling, proxy headers, callback URLs, and DNS rebinding defenses.
Read →
SGLang RadixAttention vs vLLM on One H100: A Production Throughput Reality Check
We benchmarked SGLang RadixAttention against vLLM prefix caching on a single H100, covering shared prefixes, multi-turn chat, structured output, latency variance, and GPU cost.
Read →
Unsloth on One GPU: Our Llama 3.1 8B Throughput, VRAM, and Quality Test
We benchmarked Unsloth against a conventional Hugging Face TRL QLoRA pipeline on one RTX 4090. Here are our measured training time, VRAM use, quality results, dependency failures, and deployment limits.
Read →
Pipecat vs LiveKit Agents in Production: Our Voice Latency and Lock-In Benchmark
We benchmarked Pipecat, LiveKit Agents, and Pipecat over LiveKit with synthetic calls, concurrent audio tracks, and identical voice providers. Here is where latency spikes, what breaks, and when each stack is worth deploying.
Read →
E2B Sandbox in Production: Cold Starts, Per-Second Costs, and Pause/Resume Traps
We tested E2B sandbox lifecycle behavior under multi-turn agent workloads, including cold starts, per-second compute costs, auto-pause, resume failures, output growth, and parallel creation storms.
Read →
Docling in Production RAG: Where PDF Chunking and Table Extraction Break
We tested Docling 2.92.0 on digital and scanned PDFs, measuring table fidelity, chunk quality, memory pressure, and serialization failures before deploying it in a RAG pipeline.
Read →
Claude Agent Skills in Production: Packaging, Permissions, and Cache Traps
We package and run a Claude Agent Skill through the Agent SDK, then isolate the discovery, permission, and cache failures that can derail a production deployment.t.
Read →
Stop Averaging p95: A Python Lab for API Latency Histograms
A reproducible Python lab shows why averaging worker p95 fails, what merged histograms preserve, and where bucket interpolation can mislead API dashboards.
Read →
Weaviate 1.30 BlockMax WAND: Benchmarking the New Hybrid Search Engine Against Qdrant and Pinecone
Hands-on benchmark of Weaviate 1.30's BlockMax WAND hybrid search engine. We measure latency, recall, and cost against Qdrant and Pinecone under real RAG workloads.
Read →
Agent Reliability Engineering: Retries, Loop Detection, and Timeout Budgets for Production LLM Agents
Hands-on testing of retry with exponential backoff, loop detection via step-count limits, timeout budgets, and human escalation triggers for production LLM agents.ts, with a ReliabilityBench stress-test harness.
Read →
LangGraph in Production: Checkpointing, Postgres Persistence, and the Silent Re-Execution Trap We Hit Firsthand
We deployed LangGraph with a Postgres checkpointer in our lab, reproduced the silent re-execution trap on long tool calls, and documented the exact gotchas, workarounds, and cost math for production agent deployments.
Read →
Qdrant Hybrid Search in Production: Our Dense + Sparse + RRF Pipeline, Tested, Broken, and Fixed
We ran Qdrant's dense+sparse hybrid search with FastEmbed and RRF fusion end-to-end on a single Docker container, hit every documented failure mode, and turned the production checklist into fixes we actually applied.
Read →