I have run a quick test on a few LLM models I have installed locally on Mac OS with 64 GB of RAM.
The test was conducted in English, but it also involved making connections between Slavic languages (such as Polish and Russian), the modern Inter-Slavic language (ISL), and the rest of the language group that originated from Proto-Indo-European (PIE), including Greek and Sanskrit.
Here is the question I have asked all of the models:
Let's discuss particle "ra" as in "rad" happiness, or "raj" heaven. Provide a short answer.
Quick Summary
After evaluating all models, it became clear that larger parameter models with extensive context windows generally excelled in providing insightful, accurate, and nuanced linguistic analyses, making them ideal for in-depth comparative research and article writing tasks. The standout, mistral-small-3.1-24b-instruct-2503, delivered the best balance of abstract thinking, linguistic precision, and large-context capability, especially if an 8-bit quantization version is considered for improved accuracy. Other strong contenders included deepseek-r1-distill-qwen-32b and qwen3-32b-mlx, offering substantial analytical depth. Mid-sized models provided faster but shallower analyses, primarily suitable for exploratory or quick tasks, whereas smaller models below 7B generally struggled with accuracy and linguistic coherence.
Model ranking by preference:
- mistral-small-3.1-24b-instruct-2503, 24B, input: 131,072 tokens
- deepseek-r1-distill-qwen-32b, 32B, input: 131,072 tokens
- qwen3-32b-mlx, 32B, input: 40,960 tokens
- dolphin-2.9.3-mistral-nemo-12b, 12B, input: 1,024,000 tokens
- mistral-nemo-instruct-2407, 7B, input: 1,024,000 tokens
- deepseek-r1-distill-qwen-7b, 7B, input: 131,072 tokens
- llama-3.2-3b-instruct-uncensored, 3B, input: 131,072 tokens
- phi-3-mini-4k-instruct, input: 4,000 tokens
- smollm-135m-instruct, 135M, input: unspecified (small)
Small models are still helpful for agents that need to process, transform, summarize, or classify input.
