Squeeze-Release: Iterative Pruning with Exact Structural Minimization
Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so the deployed model is no smaller than before pruning. We prese...