See how a 30B model, DFlash, vision, and 256K context fit on one 24 GB GPU—and why the fastest quant wasn't the best set...

See how a 30B model, DFlash, vision, and 256K context fit on one 24 GB GPU—and why the fastest quant wasn't the best setup. https://hackernoon.com/benchmarking-dflash-on-a-30b-model-why-tokens-per-second-can-mislead #machinelearning

Read Original

Related