Articles, one at a time.
Every piece cites primary sources and adds something you can act on: a worked example, a tested command, a documented limitation. Drafts that only summarize what already exists stay unpublished.
Google Antigravity 2.0: Can Its Browser-Driving Agents Survive Real Repo Work?
We installed Google Antigravity 2.0 in our lab and put its Agent Manager and browser-driving agents through a scripted battery of real repo tasks, including a browser-driven E2E test. Here is what worked, what broke, and whether it justifies a seat.
Read →
From Notebook to Production SLA: Running vLLM on Kubernetes with the Production Stack
A practitioner's guide to moving vLLM from a single-GPU notebook demo to a Kubernetes deployment with measured p50/p95 latency under concurrency, grounded in the vLLM production-stack release and the 2026 GLM-5.2 production SLA architecture.
Read →
Your Agent Is Not a User: Giving AI Agents Their Own OAuth Identity with Scoped, Revocable Credentials
Why sharing your users' OAuth tokens with AI agents breaks audit, least privilege, and enterprise deals — and how to run agent identity with OBO exchange on a local sandbox.ocal stack.
Read →
Sonnet 5 Pricing: Why a Cheaper Token Can Cost You More
Sonnet 5 costs a third less per token and counts your text differently. The break-even math, plus the caching and routing architecture that decides it.
Read →
Your LLM Bill Is an Architecture Problem, Not a Prompt Problem
Provider invoices bill credentials, not customers. Three LLM cost attribution patterns, a list-price model, and which control to build first.
Read →
Semantic Caching vs. Prompt Caching: Where Break-Even Sits
Two caching layers on one workload, priced from Anthropic's official cache rates. Where each pays, and the measurement that tells you which fits.
Read →
Supabase Removes logs.all on 2026-09-23: Port Your Queries First
Supabase kills the logs.all endpoint on 2026-09-23 and the replacement speaks ClickHouse SQL only. The inventory scan and query rewrites to ship first.
Read →
OpenAI Kills the Videos API With No Successor: A 4-Week Exit Plan
The Videos API and every sora-2 model shut down 2026-09-24 with an empty replacement column. The inventory scan and exit matrix to decide before the date.
Read →
Your Coding Agent's Allowlist Is Not a Sandbox
An allowlist approves names, not behaviors. Why environment poisoning turns a trusted command into code execution, and a config auditor that finds it.
Read →
The Agent Spend Cap That Admits It Can Be Exceeded
Claude's Managed Agents budget documents that it can land past its own cap. The production architecture that actually bounds an autonomous agent's bill.
Read →
Stop Parsing Raw JSON: Type-Safe LLM Output Pipelines with BAML
Why json.loads() on model output is a production liability, and how BAML's contract-first schemas and schema-aligned parsing remove a whole failure class.
Read →
Multi-Tenant LLM Gateway: Enforcing Virtual Key Budgets & Quotas
How to give every customer of your AI product a hard spend cap and rate limit with LiteLLM virtual keys, before one tenant's bug becomes your bill.
Read →