Today, I got Gemma4:12b-it-qat running at a good 70 tokens per second on the Lenovo, Nvidia 4070, llama.cpp, with multi ...

Today, I got Gemma4:12b-it-qat running at a good 70 tokens per second on the Lenovo, Nvidia 4070, llama.cpp, with multi token prediction. Got Qwen3.8:27b running at a good 6 tokens per second. Yeah that one ain't going anywhere fast. I'm downloading Gemma4:26b to see how fast it can go on this poor machine. Also, got Emacs with Emacspeak working with an Eloquence speech server, for the new Eloquence for Linux. That's how I'm using Mastodon on Linux currently.#foss #emacs #ai

Read Original

Related