Mac users: if you are running Qwen models locally, make sure you use MLX (Apple's version of PyTorch, sort of). It's the only way to get decent performance. Windows users: just buy a GeForce RTX ffs