Showing posts with label Mistral LLM. Show all posts
Showing posts with label Mistral LLM. Show all posts

mistral

If you want a straightforward choice that feels smart, responsive, and fits on a Mac with 64 GB RAM, pick Mistral-Small-3.2-24B-Instruct-2506 in q8 format. It is a compact, high-quality build that handles long documents and gives clear, helpful answers without slowing your laptop to a crawl. Think of q8 as a careful compression of the model. You keep almost all the quality, use far less memory, and get faster replies than the full uncompressed version. If you work with very large notes or codebases, this model’s long context lets it follow the thread across many pages. If you prefer extra speed over maximum accuracy, keep a smaller 12B model on hand for quick tasks, and load the 24B only when you need top quality. For most people, the 24B q8 is the best everyday default.

See the same in Interslavic:
https://snoveni.blogspot.com/2025/09/llm.html


As an Amazon Associate I earn from qualifying purchases.

Testing local LLM on multi-lingual understanding.

I have run a quick test on a few LLM models I have installed locally on Mac OS with 64 GB of RAM.

The test was conducted in English, but it also involved making connections between Slavic languages (such as Polish and Russian), the modern Inter-Slavic language (ISL), and the rest of the language group that originated from Proto-Indo-European (PIE), including Greek and Sanskrit.


Here is the question I have asked all of the models:

Let's discuss particle "ra" as in "rad" happiness, or "raj" heaven. Provide a short answer.

Quick Summary


After evaluating all models, it became clear that larger parameter models with extensive context windows generally excelled in providing insightful, accurate, and nuanced linguistic analyses, making them ideal for in-depth comparative research and article writing tasks. The standout, mistral-small-3.1-24b-instruct-2503, delivered the best balance of abstract thinking, linguistic precision, and large-context capability, especially if an 8-bit quantization version is considered for improved accuracy. Other strong contenders included deepseek-r1-distill-qwen-32b and qwen3-32b-mlx, offering substantial analytical depth. Mid-sized models provided faster but shallower analyses, primarily suitable for exploratory or quick tasks, whereas smaller models below 7B generally struggled with accuracy and linguistic coherence.

Model ranking by preference:

  1. mistral-small-3.1-24b-instruct-2503, 24B, input: 131,072 tokens
  2. deepseek-r1-distill-qwen-32b, 32B, input: 131,072 tokens
  3. qwen3-32b-mlx, 32B, input: 40,960 tokens
  4. dolphin-2.9.3-mistral-nemo-12b, 12B, input: 1,024,000 tokens
  5. mistral-nemo-instruct-2407, 7B, input: 1,024,000 tokens
  6. deepseek-r1-distill-qwen-7b, 7B, input: 131,072 tokens
  7. llama-3.2-3b-instruct-uncensored, 3B, input: 131,072 tokens
  8. phi-3-mini-4k-instruct, input: 4,000 tokens
  9. smollm-135m-instruct, 135M, input: unspecified (small)


Small models are still helpful for agents that need to process, transform, summarize, or classify input.





 



As an Amazon Associate I earn from qualifying purchases.

LM Studio with 12 and 24B local LLM models

In my LM Studio, I have been using the 12 billion and 24 billion parameter models on my relatively inexpensive Mac Studio M1, which has 64 GB of unified memory.

It also has a 1 million token input context window! That would roughly cover the entirety of J.R.R. Tolkien's "The Lord of the Rings: The Fellowship of the Ring", or approximately 400 pages of text.










The 12B model responds almost instantly and is excellent for good-quality, rapid example work.
The 24B model takes about 30 seconds to respond, but it has deep, obscure, nuanced knowledge of the world. I would have to spend 5 times more to do the same with NVidia GPUs.

Another benefit of using the "Dolphin" is that it is uncensored, which gives me direct answers to my questions without trying to "protect me" from facts like "Tiananmen Square protests of 1980", or any other enforced ideology.



As an Amazon Associate I earn from qualifying purchases.

apt quotation..