GPT-6 Sol First Impressions: Faster Builds, Smarter Defaults
Peter compares GPT-6 Sol with GPT-5.6 Sol on the same max-reasoning creative builds. See the reported time, token, quality, and reasoning-effort lessons.
Every case study, system breakdown, and field note, newest first and filterable by topic. No theory, no demos. Only systems that run in production and what it took to get them there.
This is the complete archive: all 508 posts across 54 topics, newest first, filterable by category. It holds the case studies, the model and tool comparisons, and the field notes from systems I built and run. For the curated view, where the writing sits alongside the free Claude Code skills, start at the Library instead.
The archive leans in four directions: AI Search Visibility (98), AI Agent Architecture (61), AI Tools (38), and AI Coding Agents (27). Those four account for most of what I publish, because they are where most of the client questions land.
Three places to start. The AI search visibility guide is the hub for the largest cluster and links out to every spoke in it. Grok Imagine vs Midjourney is the image model comparison, rebuilt against live leaderboard data rather than left to go stale. OutreachIQ is the longest running system breakdown here, from first prototype through to a public repo.
Peter compares GPT-6 Sol with GPT-5.6 Sol on the same max-reasoning creative builds. See the reported time, token, quality, and reasoning-effort lessons.
Matthew Berman reviews Claude Opus 5.5 with an Anthropic engineer. Learn why cost per task matters, which effort level to use, and when Fable 5.1 still fits.
Pat Simmons blind-tests Claude Opus 5.5, Fable 5.1, and GPT-6 Astra across four real builds. Compare visual quality, 3D work, gameplay, reported cost, and runtime.
Grok 4.7 closes much of the gap on coding and agentic knowledge work. Here is what xAI claims, what independent tests show, and where the value case holds.
Steve Sewell audits viral Jev demos, explains the decision model's real boundary, and shows stronger uses in email classification, browser routing, and Agent-Native tool selection.
Matthew Berman tests Jev, TypeSafe AI's fast decision model. Learn its typed questions, RLCD training, useful routing patterns, demo limits, and why zero hallucinations is not a safety guarantee.
A ten-app stress test shows where GPT-6 Astra computer use saves real effort, where direct tools win, and how to delegate visual software safely.
Eight creator tests show where GPT-6 Astra is useful now: app building, browser and iPhone control, video editing, meeting analysis, writing, and presentations.
Ten GPT-6 Astra workflows show what computer-using AI can do across robotics, Figma, presentations, research, mobile apps, Blender, Sheets, and product analysis.
A practical review of Claude Fable 5.1 across research decks, animated SVGs, an app, Blender, a product website, and a reusable prompting skill.
Ras Mic tests GPT-6 Astra on performance, security, analytics, hardware planning, UI, architecture, and computer use after 24 hours of real work.
Ben Davis, Peter Gostev, and Tom Krcha test GPT-6 Astra on a playable history of London, matcha-shop design directions, and a DEF CON puzzle with parallel agents.
Matt Wolfe tests GPT-6 Astra on SVG coding, a Three.js game, an interactive world, Blender, and Unreal Engine. Here is what the demos prove, what they do not, and how to evaluate Astra safely.
Arena AI tests Claude Fable 5.1 against Fable 5, GPT-5.6 Sol, Opus 5, Kimi K3, GLM 5.3, Qwen 3.8, Grok, and DeepSeek across 3D worlds, games, SVGs, and research interfaces.
Bijan Bowen tests Qwen3.8 Max across C++, CAD, games, writing, and multimodal coding. See what worked, what failed, what $32 bought, and its open-weight status.
AI Search pushed Claude Opus 5 through browser software, 3D reconstruction, financial video, Blender, music, vision, medical images, and research. Here is what the tests actually prove, where the model failed, and when its cost is justified.
Every post here is about a system that actually shipped. Book a free call and let's talk about what could ship for you.
Book Free 30-min CallAdd your name and email to unlock recording.
Thanks for your voice message. I listen to every one myself and I will get back to you within one business day.