Large-scale LLM inference engine
-
Updated
Sep 11, 2026 - Python
Large-scale LLM inference engine
Foundation model benchmarking tool. Run any model on any AWS platform and benchmark for performance across instance type and serving stack options.
Community maintained hardware plugin for vLLM on AWS Neuron
This Guidance demonstrates how to deploy a machine learning inference architecture on Amazon Elastic Kubernetes Service (Amazon EKS). It addresses the basic implementation requirements as well as ways you can pack thousands of unique PyTorch deep learning (DL) models into a scalable architecture and evaluate performance
CMP314 Optimizing NLP models with Amazon EC2 Inf1 instances in Amazon Sagemaker
This repository provides an easy hands-on way to get started with AWS Inferentia. A demonstration of this hands-on can be seen in the AWS Innovate 2023 - AIML Edition session.
Terraform blueprints for running AI/ML workloads on Amazon ECS - vLLM, Triton, and TEI on GPU, Inferentia2, and Fargate
Scalable multimodal AI system combining FSDP, RLHF, and Inferentia optimization for customer insights generation.
Production LLM pipeline on AWS Trainium and Inferentia: LoRA fine-tune Llama 3.1 8B on a trn1.2xlarge, ship the adapter through S3, serve it with vLLM on an inf2.xlarge, and measure everything (TTFT/TPOT percentiles, tokens/s, MFU, goodput at SLO) with compile costs included and failures recorded as receipts.
Deploy Large Models on AWS Inferentia (Inf2) instances.
End-to-end solution for cold-start recommendations using vLLM, DeepSeek Llama (8B & 70B), and FAISS on AWS Trainium (Trn1) with the Neuron SDK and NeuronX Distributed. Includes LLM-based interest expansion, embedding comparisons (T5 & SentenceTransformers), and scalable retrieval workflows.
Pseudo-spectral direct numerical simulation (DNS) of the Taylor-Green vortex on AWS Neuron: NKI kernels on Inferentia2 and Trainium1, 3-D FFTs as matmuls, all-to-all collectives inside the kernel and a libnrt C driver, in fp32 up to 512^3, checked against an fp64 oracle and Trainium1 references. Sample code for HPC and CFD engineers.
To associate your repository with the inferentia topic, visit your repo's landing page and select "manage topics."