Mac users: if you are running Qwen models locally, make sure you use MLX (Apple's version of PyTorch, sort of). It's the only way to get decent performance.
Windows users: just buy a GeForce RTX ffs
stevibe@stevibeUpdate: Qwen3.5:9b head-to-head (MLX, Apple Silicon optimized) Mac Studio M2 Ultra: 89.74 tok/s Mac Mini M4: 20.82 tok/s MLX basically doubles the speed on both machines.






