Jev + Claude Code: A Faster Agentic Coding Review Loop
Ray Amjad explores Jev as a fast decision layer for Claude Code and Codex: skill suggestions, browser checks, comment triage, code review, and a measured build-test-fix loop.
Every case study, system breakdown, and field note, newest first and filterable by topic. No theory, no demos. Only systems that run in production and what it took to get them there.
This is the complete archive: all 508 posts across 54 topics, newest first, filterable by category. It holds the case studies, the model and tool comparisons, and the field notes from systems I built and run. For the curated view, where the writing sits alongside the free Claude Code skills, start at the Library instead.
The archive leans in four directions: AI Search Visibility (98), AI Agent Architecture (61), AI Tools (38), and AI Coding Agents (27). Those four account for most of what I publish, because they are where most of the client questions land.
Three places to start. The AI search visibility guide is the hub for the largest cluster and links out to every spoke in it. Grok Imagine vs Midjourney is the image model comparison, rebuilt against live leaderboard data rather than left to go stale. OutreachIQ is the longest running system breakdown here, from first prototype through to a public repo.
Ray Amjad explores Jev as a fast decision layer for Claude Code and Codex: skill suggestions, browser checks, comment triage, code review, and a measured build-test-fix loop.
Dubibubi reports up to 91.75% lower Codex usage and demonstrates a 33.3% reduction. Here are the eleven rules, what they change, and how to test them fairly.
Andrew Warner and Matt Pocock explain how Grill Me, specs, tickets, implementation, and review make coding-agent work portable and reliable.
Greg Isenberg's sponsored Claude Code workflow rebuilt with repo context, Plan Mode, tickets, preview, layered review, routines, worktrees, and permission tiers.
Matt Wolfe built a 3D roguelite with Claude Code, then refined it with Codex. The useful lesson is where AI compressed implementation and where design, testing, balance, and judgment still dominated.
Muse Code and Muse Spark 1.2 are exceptionally fast and cheap. Theo's test reveals the data tradeoff, rate limits, strong PR triage, and weak long tasks.
Claude Code creator Boris Cherny explains why stronger models need less legacy scaffolding. Here is the eval-led delete-and-restore method, prompt audit, verification contract, and security caveat.
Theo traced T3 Code's idle GPU load to infinite status animations interacting with a fixed noise overlay. Here is why Fable and GPT-5.6 guessed wrong, what the final patch changed, and a better AI-assisted debugging workflow.
Running Claude Code on free model tiers via OmniRoute: what is genuinely free, what leaves your laptop, which routes raise terms concerns, and a safer setup.
Theo examines Linus Torvalds' defense of AI-assisted Linux kernel review. Here is the primary-source context, what Sashiko's 53.6% result does and does not prove, and the accountable workflow open-source teams can copy.
Theo argues that cheap AI code should generate debuggers, test harnesses, stress rigs, and disposable experiments around the code that matters. Here is the risk-based workflow that makes the idea safe and useful.
OpenAI's Codex plugin adds read-only reviews, adversarial audits, rescue tasks, and session handoffs to Claude Code. Setup plus a practical two-agent workflow.
Andrew Warner and Bryan McAnulty review seven GPT-5.6 Sol tests across games, dashboards, video, browser automation, writing, skills, image prompting, and Vision Pro.
Riley Brown tests Grok 4.5 inside Cursor on a research canvas, personal website, Swift voice app, Design Mode edits, and an Excalidraw-style Convex application.
Theo used GPT-5.6 across 67 projects and an estimated $180K-$240K of inference. Here is the organized ledger: production work, native rewrites, Rust experiments, browser automation, failures, and the practices worth copying.
Marcin AI tests GPT-5.6 on a company website, voice-controlled chessboard, and native iOS alarm app. The chessboard reveals what agentic building still needs: speech input, legal move validation, testing, and human correction.
OpenAI launched GPT-5.6 for Codex and ChatGPT Work. Here is the builder read on Sol, Terra, Luna, max and ultra reasoning, Programmatic Tool Calling, subagents, and the new Codex work operating system.
Grok 4.5 for coding and agentic work: benchmarks, pricing, token efficiency, Grok Build, Cursor support, EU availability, and when it beats Fable, Opus or GLM.
Theo argues the Fable 5 backlash is mostly wrong. Here is the practical builder version: Fable was not simply nerfed, but routing, safety classifiers, usage limits, and cost discipline now matter more.
OpenAI Codex Record & Replay lets you show Codex a repeated workflow and turn it into a reusable skill. Here is what it changes, where it works, and what builders should record first.
GLM 5.2 can be routed into Claude Code through Z.AI. Here is what changed, the config to use, where it may beat Opus on cost, and what builders should test first.
Anthropic studied 400,000 Claude Code sessions and found that domain expertise matters more than traditional coding experience. Here is what builders and small teams should learn.
A practical look at Claude Fable 5 build demos, from Lovable-style app builders and 3D games to custom productivity tools, plus the prompting patterns that make expensive agents worth using.
Codex is now Windows-native with PowerShell and a real sandbox. What changes for sysadmins, developers, security labs, and anyone letting an agent touch a PC.
Every post here is about a system that actually shipped. Book a free call and let's talk about what could ship for you.
Book Free 30-min CallAdd your name and email to unlock recording.
Thanks for your voice message. I listen to every one myself and I will get back to you within one business day.