Ollama 0.31 introduces a substantial speed boost for Gemma 4, with Apple Silicon users seeing nearly 90% faster token generation in coding-agent benchmarks. The performance gain comes from a multi-token prediction method that proposes and verifies several tokens at once, with the draft length auto-tuned in real time and no change to model outputs. An optimized matrix multiplication kernel contributed to MLX further accelerates batch verification, benefiting both Gemma 4 and other models. View article on AlternativeTo »More about Ollama | Ollama Alternatives
Related
xAI introduces Grok Voice Think Fast 2.0 with major accuracy and conversational gains
xAI has unveiled Grok Voice Think Fast 2.0, its latest voice AI model focused on delivering enhanced transcription accuracy and smarter conversational ability. This release introdu...
Gleam 1.18.0 adds advanced record field support and boosts JS performance
Gleam 1.18.0 introduces major upgrades to its language server, now supporting advanced go-to-definition, rename, and referencing features for record fields across modules. Develope...
Krita 5.3.3 brings bug fixes, Android supporter options, and feature changes
Krita 5.3.3 and 6.0.3 are now available, focusing on bug fixes and overall stability improvements. The Android app adds in-app donation and supporter benefits, including direct bun...