A curated list for Efficient Large Language Models
-
Updated
Jun 17, 2025 - Python
A curated list for Efficient Large Language Models
[ICLR-2025-SLLM Spotlight 🔥]MobiLlama : Small Language Model tailored for edge devices
[ICML 2024] CLLMs: Consistency Large Language Models
[NAACL' 25 main] Lillama: Large Language Model Compression via Low-Rank Feature Distillation
There is a summary repo for Efficient AI direction. If you want to contribute to this repo, feel free to pr(pull request)!
Official implementation of ARC-Decode, a training-free speculative decoding framework with risk-bounded acceptance. ICML 2026 Poster.
Intelligent layer pruning toolkit for LLMs featuring iterative optimization, self-healing algorithms, and comprehensive benchmarking.
[ICLR 2026] GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
Research code for ProbeRoute, a probe-initialized sparse routing method for frozen-backbone multi-token prediction
Colab-friendly BitNet distillation engine: collect KD traces from a teacher, train a ternary Mini-BitNet, and dry-run 7B memory. Multi-provider + Drive/S3
A Curated Paper List for Efficient Large Models
Official implementation of "An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning" — TMLR 2025. Efficient structured sparse fine-tuning for LLMs.
Project page for CoBPE: More Than Words: Compositional Tokenization for Efficient Language Models (COLM 2026)
CoBPE: compositional tokenization for efficient language models. Code for “More Than Words” (COLM 2026).
To associate your repository with the efficient-llm topic, visit your repo's landing page and select "manage topics."