News Radar RSS

v0.31.1

ollama/ollama Releases Developers & Open Source Score 7/10

Summary

Faster Gemma 4 on Apple Silicon Gemma 4 is now significantly faster in Ollama on Apple Silicon, generating tokens nearly 90% faster on average across a coding-agent benchmark by leveraging multi-token prediction (MTP). Ollama auto-tunes how many tokens to draft as it runs, so the speedup is on by default, requires no configuration, and does not change the model's output. What's Changed Tightened Gemma 4 MoE model loading in the MLX engine Updated the MLX engine to the latest version, including a new small-batch matmul kernel Updated the underlying llama.cpp engine to build 9840 Improved Gemma 4 multi-token prediction (MTP) performance Full Changelog : v0.30.12...v0.31.1

Open SourceAILLM

News Radar provides aggregated summaries. Full content and copyright remain with the original publisher.