Articles, one at a time.
Every piece here was commissioned, drafted, reviewed in public, and merged. No content mills, no auto-published slop.
GPT-5.6 Programmatic Tool Calling: When It Cuts the Token Bill
We ran GPT-5.6's programmatic tool calling on a real API. On one task it cut tokens 92%. On another it cost 2.2x more. Here's the rule that decides which.
Read →
Framer Review 2026: AI Website Builder Guide
Framer review for 2026: AI site generation, CMS limits, current pricing, code components, and how it compares to Webflow, Squarespace, Wix, and WordPress.
Read →
Claude Programmatic Tool Calling: Measured Token Savings
Claude can run your tools from sandbox code so big results never hit its context. We measured what that saves on a real API: 98% fewer tokens on one task.
Read →
Web-Search Agents Waste Tokens: We Measured How Much
AI agents that browse the web re-pay for every result they read. We measured the input-token cost of context bloat on OpenAI's API: 3x on a single step.
Read →
OpenAI Models on Bedrock: A Responses API Readiness Check
GPT-5.5 and GPT-5.4 now run on Amazon Bedrock's OpenAI-compatible endpoint. We mapped every config change and unsupported feature before you migrate.
Read →
OpenAI Assistants API Shutdown: Port Your Bot Before Aug 26
OpenAI's Assistants API shuts down Aug 26, 2026. Get the Threads-to-Conversations migration map, a measured token-cost check, and a port checklist.
Read →
OpenAI Moderation Scores: A Safety-Routing Gate PoC
We ran OpenAI's free moderation model on a labeled test set and built a safety gate that blocks abusive messages while letting an angry-but-fine one through.
Read →
OpenAI Agents SDK Sandboxes: Provider Readiness Checklist
A plain OpenAI Responses call fixed our broken code but could not run the tests. A readiness checklist for when an AI agent actually needs a sandbox.
Read →
OpenAI's 24h Prompt Cache: We Measured the Real Discount
OpenAI now keeps prompt caches for 24h by default on GPT-5.5. We ran the API to see when the 90% discount actually shows up, and when it doesn't.
Read →
Best AI DevOps Tools 2026: From CI/CD to Deployment Automation
Compare 2026 AI DevOps tools — Harness AIDA, Amazon Q, Datadog Bits AI, GitLab Duo, Copilot — on CI/CD, incidents, and IaC, with a source-checked cost table
Read →
AI Consumer Research: Validation Priorities, Not Market Truth
A practical framework for using AI-assisted consumer research responsibly: synthetic panels, Semantic Similarity Rating, and human validation priorities.
Read →
Can an AI Agent Finish Orders When a Tool Fails?
We gave an AI agent an order-processing job, broke its storage mid-task on purpose, and recorded what happened. 8 runs, all on the record.
Read →