Apple Silicon LLM inference
Apple Silicon LLM inference Articles
Browse 4 articles about Apple Silicon LLM inference.

Parallel Constrained Decoding: 7x Faster JSON on Apple Silicon
Parallel constrained decoding cuts structured JSON extraction latency up to 7x on Apple Silicon by scoring schema fields at once, not token by token.
parallel constrained decodingMLX structured outputApple Silicon LLM inference

Parallel Constrained Decoding: Faster Structured JSON on Apple Silicon
An MLX engine for Apple Silicon evaluates JSON schema fields in parallel, cutting structured extraction latency 5.6x to 7x with guaranteed valid syntax.
parallel constrained decodingMLX Apple Silicon inferencestructured JSON generation

Edge0: Run a 35B Parameter Model in Under 3GB of Memory
Edge0 streams MoE experts from disk to run Qwen 3.5 35B-A3B in under 3GB RAM. Here's how the architecture works and how to install it.
Edge0 inferencerun 35B model locallyQwen 3.5 MoE

How to Run Maple-Preview Locally on a Mac Mini M4
A guide to running DeepGrove's Maple-Preview 20B-A1B ternary model on a Mac mini M4, covering hardware needs and real-world speed.
run Maple-Preview locallyMac mini M4 LLMternary weight inference