GLM-5.2 at Home: What 0xSero's Local Inference Bet Teaches Builders
David Ondrej interviews 0xSero about GLM-5.2, custom compression, LM Studio, rented GPUs, local tokens, and why open-weight models need better distribution and tooling.
Every case study, system breakdown, and field note, newest first and filterable by topic. No theory, no demos. Only systems that run in production and what it took to get them there.
This is the complete archive: all 347 posts across 40 topics, newest first, filterable by category. It holds the case studies, the model and tool comparisons, and the field notes from systems I built and run. For the curated view, where the writing sits alongside the free Claude Code skills, start at the Library instead.
The archive leans in four directions: AI Search Visibility (73), AI Agent Architecture (52), AI Tools (27), and AI Coding Agents (21). Those four account for most of what I publish, because they are where most of the client questions land.
Three places to start. The AI search visibility guide is the hub for the largest cluster and links out to every spoke in it. Grok Imagine vs Midjourney is the image model comparison, rebuilt against live leaderboard data rather than left to go stale. OutreachIQ is the longest running system breakdown here, from first prototype through to a public repo.
David Ondrej interviews 0xSero about GLM-5.2, custom compression, LM Studio, rented GPUs, local tokens, and why open-weight models need better distribution and tooling.
A practical comparison of cloud GPUs, hosted AI APIs, and home AI hardware for local models: cost, privacy, latency, maintenance, electricity, and when a hybrid setup wins.
A practical local AI hardware guide with prices and buy links for Ollama, LM Studio, Qwen, Gemma, Llama, DeepSeek, GLM, Mac mini, Mac Studio, RTX PCs, DGX Spark, and cloud GPUs.
NVIDIA Nemotron 3 Ultra is a 550B open MoE model for long-running agents, while Nemotron 3 Super is the 120B efficient workhorse. Here is what builders should know about architecture, licensing, deployment, and local experiments.
Google's own AI-search guidance is much less mystical than most SEO chatter. Here is what the docs actually point toward, and what they do not support.
AI can speed up web production, but it can also flatten brands into the same safe average. Here is the stack I would use to keep AI-assisted websites distinctive enough to survive AI-mediated discovery.
Higgsfield AI shows how to make a cinematic product commercial with GPT Image 2, Soul Cinema, a Claude shotlist skill, and Seedance 2.0 4K. Here is the practical production workflow.
Austin Marchese breaks down the B.U.I.L.D. framework for making Claude Code improve over time. Here is the practical version: knowledge base, bulk ingest, data inflow, review loops, and routines without system drift.
Dan Kieft previews Seedance 2.5 and tests Seedance 2.0 4K for AI filmmaking. Here is what is confirmed, what is still pre-release, and how builders should test 4K video without burning credits.
AI Samson compares NVIDIA Cosmos 3 and Ideogram 4 against GPT Image 2 and Nano Banana Pro. Here is the practical builder guide to when free open-weight image models are enough, and when paid models still win.
A practical AI weekly roundup covering GPT-5.6 limited access, Claude Tag in Slack, NVIDIA BioNeMo, Codex mobile, Gemini study notebooks, Notion agents, Figma Motion, open models, and what builders should test.
Pat Simmons shows three ways to reduce dependence on gated frontier models: local Ollama, free NVIDIA NIM endpoints, and cheap OpenRouter model routing. Here is the practical builder version.
Wes Roth walks through Hermes Agent with Stripe payments, Nous Portal, OpenAI Codex, NVIDIA NemoClaw, and Nemotron 3 Ultra. Here is the practical builder version: agents can run business workflows, but money movement needs hard approval gates.
A practical local AI starter guide inspired by Alex Finn: what local AI is, why it matters, what hardware and software to start with, and which private workflows are worth testing first.
If AI agents increasingly compare, shortlist, and even act on behalf of users, service sites need stronger trust, clarity, and proof layers. Here is what that changes.
Bing's June 2026 AI visibility upgrades made citation tracking much more concrete. Here is why Citation Share matters even if Bing is not your main lead channel.
Greg Isenberg interviews Matt Van Horn about Last 30 Days, a Claude Code skill that researches X, Reddit, YouTube, Hacker News, Polymarket, GitHub, and the web so prompts start with current intelligence.
OpenAI officially previewed GPT-5.6 Sol, Terra, and Luna. Here is what is worth using, what is still gated, and what builders should test before switching.
OpenAI previewed GPT-5.6 Sol, Terra, and Luna, but access is limited to trusted API and Codex partners. Here is what is official, what the reaction videos get right, and what builders should do next.
A clean homepage is not enough for AI-mediated discovery. Here is how to think in proof packs: linked sets of service, system, founder, and third-party evidence that help AI systems trust your brand claims.
Nate Herk and Suvam Khadka break down a cold email framework that generated $500K+ in sales opportunities. Here is the practical AI outreach version: zero-risk offers, niche lead sources, personalization, and compliance.
Should you hire an AI consultant or build automation in-house? A practical decision framework covering cost, speed, risk, control, and when each option wins.
What does AI automation actually cost? Real budget ranges for quick builds, full systems, and consulting, plus the factors that move the price up or down.
Nate Herk says four Claude Code upgrades helped him make more money: stress-test ideas, verify outputs, manage context, and use sub-agents. Here is the practical JQ AI SYSTEMS version.
Every post here is about a system that actually shipped. Book a free call and let's talk about what could ship for you.
Book Free 30-min CallAdd your name and email to unlock recording.
Thanks for your voice message. I listen to every one myself and I will get back to you within one business day.