From-scratch Rust+CUDA LLM inference engine for RTX 5090 (sm_120a) — NVFP4, MoE, MTP speculative decoding, tuned against measured hardware limits
Related
alphaparkinc/genpark-multimodal-generative-art-prompt-synthesis-engine-skill: Multimodal generative art prompt synthesis & diffusion renderer (SeaArt style)
Multimodal generative art prompt synthesis & diffusion renderer (SeaArt style)
Greninja9257/LabLLM: A native macOS lab for teaching tiny language models to think — build the architecture, train the weights, and watch a small LLM emerge from scratch, locally on Apple Silicon with custom data, tokenizers, checkpoints, and MLX acceleration.
A native macOS lab for teaching tiny language models to think — build the architecture, train the weights, and watch a small LLM emerge from scratch, locally on Apple Silicon with ...
cneuralnetwork/oplogs: Local-first, open-source experiment tracking for machine learning and agents
Local-first, open-source experiment tracking for machine learning and agents