Fully Offline Local AI assistant is hungry for GPU memoryRunning Qwen3.6 27B Q8_0 with 256k context in reasoning mode loads around 50GB of the GPU memory and gives around 64 tokens/s for prompt+generation and that is quite good for a local model with that much context.Originally published on My Tech Blog:https://minox.cosmichive.com/experiences-from-setting-up-fully-offline-local-only-ai-assisted-workstation/#ai #aiagents #linux #llm
Related
Wird nun doch vermehrt auf die robots.txt geachtet? Wäre wünschenswert.KI-Training und #Urheberrecht: Heftiger Streit üb...
Wird nun doch vermehrt auf die robots.txt geachtet? Wäre wünschenswert.KI-Training und #Urheberrecht: Heftiger Streit über Daten und Lizenzen | heise online https://www.heise.de/ne...
The commercial success of #AI products depends entirely upon the ability of its creators to incite panic and #FOMO among...
The commercial success of #AI products depends entirely upon the ability of its creators to incite panic and #FOMO among their investors. And nobody plays that game better than thi...
ははっ、JですかApple、Mac向けにmacOS 27 Golden GateのBeta版で先行して修正された20件以上の脆弱性を修正した「macOS Tahoe 26.6.2」をリリース。 https://applech2.com/ar...
ははっ、JですかApple、Mac向けにmacOS 27 Golden GateのBeta版で先行して修正された20件以上の脆弱性を修正した「macOS Tahoe 26.6.2」をリリース。 https://applech2.com/archives/20260818-macos-tahoe-26-6-2-security-update.html#Appl...