Model Compression: Make Your Machine Learning Models Lighter and Faster
A deep dive into pruning, quantization, distillation, and other techniques to make your neural networks more efficient and easier to deploy.
Writing
Notes on machine learning, deep learning, data science, and mathematics.
A deep dive into pruning, quantization, distillation, and other techniques to make your neural networks more efficient and easier to deploy.
How mathematical structure can make neural-network layers smaller and faster, from Toeplitz matrices to Butterfly and Monarch layers.
A concise proof demonstrating that there does not exist a set that contains all sets!
Estimating the number of wasted parking spaces caused by random parking.
Derivation of the formula for the volume of the n-ball.