DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
-
Updated
Mar 13, 2025 - Python
DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
From-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path commented for the why. Educational, not a llama.cpp replacement.
Dashboard for AI Studio, Open Source Continuous Inference | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro
LLM benchmarking, GPU workload orchestration backend server | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro
Sub-real-time MiniMax H3 on Hopper: 13.506 s for a 14.375 s 768p video with audio on 8xH100. Four runtime patches removing 9.72 s of non-model overhead from FastVideo.
Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines
A Flexible and High-Performance Inference Serving Engine for Open Diffusion Language Models
Rent ready-to-use cloud GPUs in seconds. Lium CLI makes it easy to launch, manage, and scale GPU compute directly from your terminal. Fast, cost-optimized, and built for AI & ML developers.
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
CLI tool to check Oracle Cloud compute shape availability - find capacity across regions
Bunch of explanations and tutorials around confidential computing
PyTorch DDP and multi-GPU LLM training — from a minimal distributed example to NanoGPT speedruns, Muon, H100 profiling, and modded-nanogpt. FBA LAB https://bubblnet.com
Demo for installing ComfyUI on Azure VM powered by Nvidia H100 to run Text to Image and to Video models like Z-Image, Qwen and Wan
From-scratch 125M LLM: 300K steps, 57.8B-token corpus, 282K tok/s on 2x H100
In the recent competition, we were challenged to finetune a model that can convert a LaTeX expressions into Python code effectively. My team, which I led, secured 6th place overall.
High-performance Triton kernels for NVIDIA H100. Implements fused FP8 LayerNorm, tiled FlashAttention, and SRAM-optimized memory primitives for Hopper architecture.
Collection of Flash attention 3 precompiled wheels for direct install on Hopper series NVIDIA GPU (H100, H200 sm_90a)
To associate your repository with the h100 topic, visit your repo's landing page and select "manage topics."