Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation preconditioner and more)
-
Updated
May 30, 2026 - Python
Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation preconditioner and more)
A performance-optimized Muon optimizer implementation for PyTorch
High-performance CUDA implementation of Muon optimizer for LLM training. Features Newton-Schulz polar decomposition, cuBLAS acceleration, and transpose optimization for 8x FLOP savings on transformer FFN layers. Benchmarked on NVIDIA A100 with Llama 3.1 8B architectures (4096×11008 weights).
The ultimate learning resource for the Muon optimizer - Newton-Schulz orthogonalization, theory, code examples, and production guides
MuonLab — Interactive Optimizer Geometry & Spectral Descent Laboratory. 6 modules: norm-ball steepest descent, Newton-Schulz orthogonalizer, update anatomy, NS-vs-SVD cost, modular duality, nanoGPT speedrun timeline.
A high-performance, VRAM-efficient PyTorch optimizer for LLMs. Combines Newton-Schulz (Muon-style) orthogonalization for 2D weights with AdamW for 1D vectors. Prevents singular value collapse and is natively compatible with FSDP & Hugging Face.
Add a description, image, and links to the newton-schulz topic page so that developers can more easily learn about it.
To associate your repository with the newton-schulz topic, visit your repo's landing page and select "manage topics."