Skip to content
Effloow
~/articles · 105 PIECES

Articles, one at a time.

Every piece cites primary sources and adds something you can act on: a worked example, a tested command, a documented limitation. Drafts that only summarize what already exists stay unpublished.
AI DEVELOPMENT 2026-08-23 ·Effloow Editorial

Beyond Naive Summarization: Context Compaction with Temporal Memory

Why long-running AI assistants bleed money on repeated history, and the memory architecture (compaction, caching, extraction) that stops it.
Read →
AUTOMATION 2026-08-22 ·Effloow Editorial

Why Agent While-Loops Fail in Production: Build Durable Workflows

Agent while-loops lose state on every crash, timeout, and rate limit. How durable execution with Inngest AgentKit or Temporal fixes it, with code blueprints.
Read →
AI INFRASTRUCTURE 2026-08-21 ·Effloow Editorial

RouteLLM in Production: Dynamic Cascades That Cut LLM Spend up to 85%

Why AI products bleed money on frontier models, and the three-layer architecture (pruning, caching, hybrid routing) that fixes cost without killing quality.
Read →
AI INFRASTRUCTURE 2026-08-19 ·Effloow Editorial

gpt-realtime Retires Jan 2027: One API Call Audits Your Stack

OpenAI retires nine voice model IDs on 2027-01-20. We probed the live API: no warning headers, but a shutdown_date field flags 52 of 126 models.
Read →
AI INFRASTRUCTURE 2026-08-17 ·Effloow Editorial

Adding One Tool to Your Agent Wiped the Whole Prompt Cache

Effloow Lab ran 17 OpenAI API calls. Appending, deleting, reordering or rewording a single tool zeroed the prompt cache every time. One setting avoided it.
Read →
AI INFRASTRUCTURE 2026-08-14 ·Effloow Editorial

The gpt-5 Alias Points at a Model OpenAI Deletes Dec 11

We asked the live OpenAI API which model answers when you write gpt-5. It is the exact snapshot being deleted on December 11, 2026.
Read →
AI DEVELOPMENT 2026-08-10 ·Effloow Editorial

Your Prompt Tooling Has a Deadline Your Monitoring Can't See

Two vendors are switching off prompt and eval tooling. We probed the live APIs, found no warning headers, and built a scanner that dates every hit.
Read →
AI INFRASTRUCTURE 2026-08-03 ·Effloow Editorial

Instrument Multi-Step Agents with OpenTelemetry GenAI Conventions: Tracing That Survives Production

A practical guide to instrumenting multi-step LLM agents with the OpenTelemetry GenAI semantic conventions, so a 2 AM failure becomes a readable trace instead of a support ticket. Includes a runnable local demo with no paid API keys.
Read →
AI INFRASTRUCTURE 2026-08-02 ·Effloow Editorial

Your Next MCP Server Could Be the Breach: A Supply-Chain Vetting and Quarantine Pipeline for Third-Party Tool Servers

A practical vetting and quarantine pipeline for third-party MCP servers: scan install scripts, verify provenance, and run candidates in a secret-free sandbox before they ever touch your agent stack.
Read →
AI INFRASTRUCTURE 2026-08-01 ·Effloow Editorial

MCP Tool Poisoning Is Now an OWASP-Listed Attack: Build a Description-Hygiene and Provenance Auditor for Your MCP Servers

MCP tool poisoning — hidden instructions in tool descriptions that hijack agents — is now an OWASP-listed attack. Build a fully local description-hygiene and provenance auditor you can hand to enterprise security reviewers.
Read →
AI INFRASTRUCTURE 2026-07-29 ·Effloow Editorial

OpenAI Spend Limits Return 429: Your Retry Logic Will Make It Worse

OpenAI hard spend limits reject requests with a 429. Lab evidence on why default client retries treat it as a temporary blip, and the one-line fix.
Read →
AI INFRASTRUCTURE 2026-07-27 ·Effloow Editorial

Your AI Gateway Blocked the Request. Did You Still Pay?

We measured what OpenRouter returns when a request is refused, and whether refused calls still cost money. Ten blocked calls moved the bill by $0.00.
Read →

Weekly field notes.

One short dispatch with new guides, tools, and what we tested that week.

Get weekly AI tool reviews & automation tips

Join our newsletter. No spam, unsubscribe anytime.