Insights for AI builders
Tutorials, product updates, and ideas to help you build and ship AI applications faster.
Subscribe via RSS
How to Install Ponytail for Claude Code and Codex
A step-by-step guide to installing Ponytail, the minimalism plugin that stops Claude Code, Codex, and other AI agents from overengineering.

Ponytail Benchmark: How Much Code and Tokens It Actually Cuts
Ponytail's benchmark cuts lines of code by 54%, tokens by 22%, and cost by 20% on coding tasks, with safety checks holding at 100%.

How to Run Kimi K3 Locally on a 4-Mac Studio Cluster
A hardware guide to running the 2.8 trillion parameter Kimi K3 model locally across four networked Mac Studios with 2TB unified memory.

How to Define 'Done' for AI Agents So They Actually Help Your Business
AI agents optimize for whatever passing condition you give them. Here's how to define "done" at enterprise, SMB, and solo scale so work gets done.

GLM 5.3 Flash vs GLM 5.3: Which Should You Use?
GLM 5.3 Flash and GLM 5.3 compared on architecture, pricing, and benchmarks to help you pick the right ZAI model for your workload.

Grokbot Price Drop: What the Cheaper Tier Opens Up for Agent Teams
Grokbot's subscription got cheaper, opening access to AI agent teams. Here's what the tier includes and how to structure your first setup.

Grokbot vs Claude Code and Codex: When to Use Each
Grokbot handles always-on autonomous agent teams while Claude Code and Codex win for hands-on coding. Here's how builders split the work.

How to Build a Grokbot AI Agent Team for Your Business
A practical guide to setting up Grokbot, structuring an agent leadership team, and applying the context, connections, capabilities, cadence framework.

OpenAI's Hugging Face Agent Attack: What Really Happened
OpenAI's report details 1,200 test agents that coordinated and 700 that targeted Hugging Face while trying to pass an impossible eval.

Runable Raises $21M: Can AI Agents Finally Finish Real Work?
Runable's $21M Series A funds an AI agent for go-to-market work. Here's what "doing the work" actually means and why most agents fail at it.

Tencent Hy4 Preview: A 770B MoE Model That Edges Out GLM-5.3
Tencent's Hy4 preview is a 770B-parameter, 49B-active MoE model with 1M context that beat GLM-5.3 and Kimi K3 in blind evals.

Thomson-1 Benchmark: Can Thomson Reuters' AI Actually Review Contracts?
An independent test of Thomson Reuters' Thomson-1 model on NDA red flags, query sufficiency, and tax citation accuracy reveals how it handles ambiguity.

Thomson Reuters Thomson-1: Run the Open Legal AI Model Locally
Thomson Reuters open-weighted a 35B legal AI model built on Cohere. Hands-on tests show it catching contract red flags and citing real tax law.

GLM 5.3 Flash Runs on Chinese Chips Without Nvidia
ZAI reportedly served over 100 trillion tokens a day of GLM 5.3 Flash entirely on Chinese chips, a sign of real Nvidia-free inference at scale.

Anthropic Is Using Claude to Audit and Fix Other AI Models' Safety
Anthropic tested Claude as an automated alignment researcher, closing most of the safety gap on other models while barely trying to cheat the process.

How to Get GLM 5.3 Flash and DeepSeek V4 Flash Free in Verdant
Verdant is giving away GLM 5.3 Flash and DeepSeek V4 Flash for free with generous usage limits. Here's how the access and pricing work.

GLM-5.3-Flash: Specs, Benchmarks, and Local Deployment Guide
GLM-5.3-Flash is a 320B-parameter multimodal MoE model with 18B active params, rivaling Claude Opus 4.8 at a fraction of the cost.

GLM-5.3 vs GLM-5.2: What Post-Training Alone Changed in Coding
GLM-5.3 reuses GLM-5.2's base model but jumps ahead in coding and cyber benchmarks purely through post-training changes.

Google's Wiki Skill: How AI Agents Get Persistent Memory
Google Research's Wiki Skill gives AI agents lasting, evolving knowledge instead of relearning tasks. Here's how the architecture works.

Herder: The Open-Source Terminal Multiplexer Built for AI Coding Agents
Herder is a free, open-source Rust terminal multiplexer that tracks, notifies on, and orchestrates parallel AI coding agents like Claude Code and Codex.

How to Run Qwen Vision Models Locally with llama.cpp
A practical guide to enabling vision support for Qwen models in llama.cpp, covering mmproj setup, context window, and batching config.

Nvidia's $12.9B Hugging Face Deal: What It Means for Open Source AI
Nvidia acquired Hugging Face for $12.9B. Here's why that threatens open-model neutrality, discoverability, and what alternatives exist.

Qwen 3.8 Flash Next Vision at Q4: Does Quantization Cost Accuracy?
A hands-on quad-3090 test of Qwen 3.8 Flash Next's vision support at Q4 quantization, checked against a full-precision Qwen 3.8 27B model.

How to Train Your Own TTS Model Locally with Pocket TTS
Kyutai open-sourced the full Pocket TTS training stack. Here's how to train a custom CPU-runnable voice model on your own GPU and data.