ISTA-DASLab/lossless-lm-dev
Updated
None defined yet.
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation