Software engineer building production AI systems: agents, retrieval pipelines and the data platforms underneath them. I learn in public by writing long-form, first-principles books and guides on LLMs, GPUs and ML systems.
π ishwarj.com: what you will find there
Everything on the site is written from first principles, in simple English, with code you can run and numbers you can measure. Dark mode by default.
| Book | What it covers |
|---|---|
| GPU Programming | Predict how fast a kernel can go, measure it, and close the gap. |
| RAG: From First Principles | Retrieval-augmented generation built and measured step by step: embeddings, chunking, hybrid search, evaluation. |
| LLM from Scratch | A language model end to end: tokenizer, data pipeline, architecture, pretraining, fine-tuning. |
| LangGraph & LangChain | A code-first deep dive into building agents with LangGraph and LangChain. |
| Designing Data Intensive Applications | Chapter-by-chapter notes on data systems, starting with the trade-offs in data systems architecture (in progress). |
Landmark papers read slowly: the paper's own lines highlighted, plain-English explanations, definitions, figures and real code that checks every claim.
- BERT (Devlin et al., 2018), in 6 parts: the big idea, architecture and input, pre-training, fine-tuning and results, ablations, impact.
- Attention, From the Ground Up (4 parts): self-attention, MQA / GQA / MLA, sliding-window and sparse attention, linear attention, Gated DeltaNet and hybrid models.
- LLM Inference (4 parts): prefill and decode, the KV cache, how vLLM serves thousands of users, speculative decoding.
- HNSW, From the Ground Up: how vector search really works.
- How Vector Search Really Works: HNSW, Filters and Qdrant
- Jev: The AI Model That Doesn't Talk. It Decides.
- π€ AI systems in production: LLM agents, RAG pipelines, vector search, evaluation.
- ποΈ Backend and data platforms: scalable APIs, event-driven systems, ETL pipelines and warehouses.
- β‘ ML systems and performance: LLM inference, attention variants, GPU programming.
- βοΈ Certified AWS Solutions Architect.
- Consumr.ai: data ingestion pipelines from the Facebook Ads and Google Ads APIs, PostgreSQL tuning, Flask APIs and a marketing-spend analytics platform.
- Cart.com: Airbyte and Apache Airflow workflows bringing Shopify data into Snowflake, with SQL stored procedures and Python ETL.
- Vision IT Labs: led the backend team; notification service on Django, Celery and Kafka; stats service on Kafka Connect and BigQuery; EKS, Jenkins and Terraform.
- Yes Lawyer: backend services with Twilio and OpenAI, a transcription system, and Azure infrastructure.
- CVS Health: large-scale healthcare data pipelines on Databricks, secure Google API access and gRPC integrations.
| Area | Tools |
|---|---|
| LLMs and NLP | PyTorch, Hugging Face Transformers, LangChain, LangGraph, RAG, vector databases (Qdrant, Weaviate, FAISS) |
| Machine learning | scikit-learn, XGBoost, LightGBM, CatBoost, TensorFlow |
| Backend | Python (Django, FastAPI, Flask), REST and gRPC, Celery |
| Data engineering | Apache Airflow, Spark, Kafka, Snowflake, Redshift, BigQuery, Databricks |
| Databases | PostgreSQL, MySQL, Redis, MongoDB |
| Cloud and DevOps | AWS, GCP, Azure, Docker, Kubernetes (EKS), Terraform, GitHub Actions, Jenkins |
Start reading at ishwarj.com Β· new books, papers and videos are added regularly.



