Fused Triton kernels for TurboQuant KV cache compression — 2-4 bit quantization with RHT rotation. Drop-in HuggingFace & vLLM integration. Up to 4.9x KV cache compression for Llama, Qwen, Mistral, and more.
Related
gargantuanfo/stable-diffusion-crack-2026: Stable Diffusion crack — run SDXL locally without GPU limits or API billing.
Stable Diffusion crack — run SDXL locally without GPU limits or API billing.
vukrosic/speculative-decoding-research: Evidence-first lab for speculative decoding, DFlash/DFlash2, block diffusion, exactness, acceptance, and inference performance
Evidence-first lab for speculative decoding, DFlash/DFlash2, block diffusion, exactness, acceptance, and inference performance