News Radar RSS

v0.32.4

ollama/ollama Releases Developers & Open Source Score 7/10

Summary

What's Changed Support Laguna on Apple GPUs via the MLX engine Quantize draft-model output heads at the requested type when creating speculative-decoding drafts. Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max). Full Changelog : v0.32.3...v0.32.4

Open SourceAILLM

News Radar provides aggregated summaries. Full content and copyright remain with the original publisher.