← All articles
Published 01 October 2026

Introduction

Local AI on Apple Silicon is now a practical, fast, and private option for many developers and power users. This guide (accurate as of October 1, 2026) walks you through a repeatable setup on a MacBook Air M2 using Ollama for model management and MLX (mlx‑lm) for Apple‑native, high‑performance inference. You’ll get runnable commands, model recommendations for an M2 machine, troubleshooting tips, and security notes so you can run Llama 3 and Mistral locally without surprises. (developer.apple.com)

What you need (short prerequisites)

If you’re short on RAM (8 GB), prefer small 3B or lightweight quantized models (Llama 3 3B or Mistral 7B in 4‑bit). If you have 16 GB or more, 7B/8B models are comfortable; 13B+ models will likely swap and be very slow on an Air M2. (smeltcore.com)

Step 1 — Install Ollama (model manager / CLI)

Ollama is the easiest way to pull and manage popular local LLMs (including Llama 3 and many Mistral variants). Install with the official script or Homebrew; the install script is the simplest:

Ollama’s quickstart docs include examples like ollama run llama3.2 and explain the pull/run workflow. (ollama.com)

Notes: - Keep Ollama updated; MLX support landed as a preview in 2026 and newer builds add more MLX-enabled models and performance improvements. If you need MLX acceleration inside Ollama, update to the current release and check the release notes. There are cases where community or custom builds are required for the experimental MLX image features — see the project notes if you hit “build with mlx” errors. (ollama.com)

Step 2 — Install MLX (mlx‑lm) — Apple’s fast runtime for M‑series

MLX (and the community package mlx‑lm) is the Apple‑native array/runtime that delivers large speedups on M‑series chips. You can either let Ollama use MLX (if your Ollama build supports it) or run MLX directly for raw performance.

Install options (choose one):

Hugging Face and mlx‑lm docs show the CLI usage (mlx_lm.generate / mlx_lm.server) and how to point at models on HF or a local cache. MLX is macOS/ARM‑only. (huggingface.co)

Step 3 — Pull and run Llama 3 (Ollama)

Ollama already hosts Llama 3 variants and makes them trivial to run.

Ollama’s docs and blog show ollama run llama3.2 examples; pick the model size that fits your RAM (3B or 8B for Air M2). If your Ollama build supports MLX backend for specific models, Ollama may choose a faster MLX runner automatically for MLX‑tagged models. (ollama.readthedocs.io)

Step 4 — Run Mistral with MLX (mlx‑lm direct)

Mistral 7B is a great fit for a MacBook Air M2 when you use quantized MLX builds (4‑bit). You can run Mistral either via Ollama (many builds available as mistral:7b) or directly with MLX for maximum throughput and lower latency.

Direct MLX server example (after installing mlx‑lm):

Hugging Face MLX docs and the mlx‑community model pages show these exact usage patterns and provide MLX‑formatted quantized models you can download. If you prefer Ollama, you can also ollama pull mistral and run through Ollama’s CLI. (huggingface.co)

Step 5 — Hybrid workflows and choosing which runtime

If your version of Ollama shows no MLX acceleration for a model, try running that same model with mlx‑lm to compare performance (some users report significant gains for certain models). (ollamaherd.com)

Performance tips (MacBook Air M2 specific)

Troubleshooting & security notes

Quick checklist (commands summary)

Conclusion

On an M2 MacBook Air you can get a practical, private local LLM environment today: Ollama for simple management and MLX (mlx‑lm) for raw Apple‑native speed. Choose quantized 3B–8B models for an Air M2, keep tools up to date, and run MLX directly when you need maximum throughput. If you hit a build or memory error, try the smaller quantized model and compare Ollama vs mlx‑lm runs — you’ll often find MLX gives the best latency on M‑series hardware. For further reading and downloads, see the Ollama quickstart, the MLX docs on Hugging Face, and the mlx‑community model pages. (ollama.readthedocs.io)

If you’d like, I can: - produce a one‑click script for installing Ollama + mlx‑lm and pulling a small Llama 3/Mistral pair tuned for 16 GB vs 8 GB, or
- help you pick the specific model tags on Hugging Face that will fit your exact MacBook Air M2 RAM configuration (tell me how much RAM your Air has).