4-bit quantization has become the de facto standard for running LLMs on consumer hardware - a 70B model drops from ~140G...

4-bit quantization has become the de facto standard for running LLMs on consumer hardware - a 70B model drops from ~140GB to ~35-40GB, and kernels like Marlin push inference close to a theoretical 4x speedup. But accuracy loss isn't uniform: general chat tasks barely budge, while coding and reasoning benchmarks have shown drops as steep as 7-14 percentage points under aggressive settings.https://psyll.com/articles/technology/ai-machine-learning/4-bit-quantization-the-real-trade-offs-explained#ai #llm

Read Original

Related

Mastodon discussion 41m ago

メキシコ!これはユグドラシルのみなさんにも教えてあげないと加賀電子、米国に営業子会社設立 脱・中国生産受け北米市場を開拓 https://www.nikkei.com/article/DGXZQOUC103RN0Q6A810C2000000...

メキシコ!これはユグドラシルのみなさんにも教えてあげないと加賀電子、米国に営業子会社設立 脱・中国生産受け北米市場を開拓 https://www.nikkei.com/article/DGXZQOUC103RN0Q6A810C2000000/#Apple #LLM #news #bot