Blog
Thoughts on AI systems, LLMs, and building software that works in production.
-
What a penetration test costs, according to three public records
The quoted average is $18,300 and traces to a 2023 page that names none of its sources. Sellers who publish put a fixed test at $4,655. The US government paid a median of $39,146 for short engagements. All three are real, and the dataset is public.
security public-records m37 -
What changed at the AI frontier: August 2026
The August edition of my monthly reference-doc diff: Google's Imagen shutdown and the migrations that are no longer model-string swaps, OpenAI's GPT-5.6 price cuts recomputed, MiniMax H3's open weights and the licence that excludes the EU, why video-understanding benchmarks are weaker than they look, and the Munich GEMA ruling against Suno.
frontier-delta ai-models deprecations llm -
What changed at the AI frontier: July 2026
A scheduled job refreshes my reference docs on AI capabilities every month and commits the diff. The July edition: OpenAI and Google retiring image and video models on short timelines, Voyage's contextualized chunk embeddings, a speech leaderboard reshuffle, Apple's 20B on-device model, and Mistral OCR 4 pricing.
frontier-delta ai-models deprecations llm -
Temperature Barely Diversifies an LLM. The Entropy Has to Enter Through the Input.
Sample the same prompt from gemini-2.5-flash ten times and you get about four distinct answers. Raising temperature from 0.2 to 2.0 barely moves that number. What does move it: seeding each call with a different persona, or showing the model its previous answers and asking for a new one. Measured on NoveltyBench, with the honest caveats about where each lever wins.
llm diversity generate-and-select ai-systems -
What it takes to ship a working demo instead of a slide deck
For a structured buyer with a real software requirement, a working demonstrator beats a slide deck. It lets both sides agree on the deliverable, and it lets the buyer verify the claims before committing.
prototyping accessibility procurement -
Claude Code Shipped Deterministic Agent Orchestration. It's Behind Two Feature Flags.
Claude Code v2.1.147 added a Workflow tool: sandboxed JavaScript that composes subagents with agent(), parallel(), pipeline(), and a real token budget. Here's what it actually does, how to enable it past the dual flag-gate, and what the empirical token cost looks like.
claude-code agent-orchestration ai-systems -
The EPD Cost Crisis: Why Small Building Materials Manufacturers Are Getting Locked Out
Environmental Product Declarations cost $17,000-$55,000 per product. For a small manufacturer on 5% margins, that's brutal. Here's where the money goes, why self-service tools don't fix it, and what just changed.
epd building-materials sustainability buy-clean -
AI Support for Premium Service Businesses: The Case for Being Available
For high-touch businesses, AI support isn't about cutting costs. It's about capturing the leads you're currently losing to voicemail.
ai customer-support premium-services
Occasional notes on AI in practice
What is changing at the AI frontier and what I am learning building with it. No spam. Unsubscribe anytime.