Showing posts with label MLX. Show all posts
Showing posts with label MLX. Show all posts

Qwen 32B

I am pleased with the performance and depth of the 32B Qwen MLX, running
locally on my Mac Studio M1 with 64GB of RAM.

9 tokens per second is not fast, but acceptable.
A 15-second wait for the first token output is very good.








As an Amazon Associate I earn from qualifying purchases.

How to get a model from HuggingFace on Mac OS

How to get a model from HuggingFace on Mac OS

This guide documents the steps needed to download HuggingFace models (especially MLX models) correctly on Mac OS.






As an Amazon Associate I earn from qualifying purchases.

mlx-lm

MLX LM is a Python package for generating text and fine-tuning large language models on Apple silicon with MLX

https://pypi.org/project/mlx-lm/#description

% pip install mlx-lm


As an Amazon Associate I earn from qualifying purchases.

apt quotation..