Popular repositories Loading
Repositories
- flex_shard Public
Flexible parameter sharding and distributed optimization utilities for PyTorch training
- tritonbench Public
Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.
- spmd_types Public
This module defines a type system for distributed training code, based off of JAX's sharding in types, but adapted for the PyTorch ecosystem.
-
-
- dist_moe Public
Dist-MoE implements Speed-of-Light distributed Mixture-of-Experts forward and backward kernels. They are highly optimized and fused comm-compute kernels for training and inference on SM100+ GPUs
-
- breakable-cuda-graphs Public
Breakable CUDA graph capture for PyTorch - capture compatible regions as CUDA graphs while running incompatible operations eagerly between them.
Top languages
Loading…
Most used topics
Loading…