Neural Foundations & Activation Math
Perceptrons, linear separability, forward pass matrix multiplication, and non-linear activations (ReLU, GELU, Sigmoid, Swish).
Backpropagation Calculus & Optimizers
Chain rule partial derivatives, computation graphs, Stochastic Gradient Descent (SGD), Momentum, AdamW, and learning rate schedules.
Deep Architectures & Vanishing Gradients
Convolution kernels, ResNet skip connections, Batch Normalization vs Layer Normalization, and vanishing gradient remedies.
Vector Embeddings & Latent Space
Word2Vec, tokenization tokenizers (BPE), cosine similarity geometry, high-dimensional projections, and semantic vector arithmetic.
Attention Mechanisms & Transformers
The Scaled Dot-Product Attention equation, Query-Key-Value (QKV) projections, Multi-Head Attention, and RoPE positional encodings.
LLM Pretraining, Perplexity & Scaling
Causal autoregressive masking, Next-Token Prediction loss, Perplexity metrics, compute optimal Chinchilla scaling laws, and dataset curation.
Fine-Tuning, LoRA & PEFT Adapters
Low-Rank Adaptation math (W = W0 + B*A), 4-bit Quantization (QLoRA NF4), Direct Preference Optimization (DPO), and RLHF.
RAG Architecture & Vector Indexing
Hierarchical chunking, HNSW vector indexing algorithms, Cross-Encoder re-ranking, and hybrid dense/sparse keyword search.
Prompt Engineering & Chain-of-Thought
Zero-Shot, Few-Shot exemplars, Chain-of-Thought (CoT), Tree-of-Thought search, system prompt sandboxing, and structured JSON generation.
Autonomous Agent Loops (ReAct)
The Reason + Act loop, tool-calling execution, reflection / self-correction loops, and stateful memory management.
Multi-Agent Swarms & Hierarchies
Hierarchical supervisor-worker agent topologies, consensus voting mechanisms, LangGraph state machines, and human-in-the-loop controls.
Production LLM Deployment & KV Cache
High-throughput inference with vLLM, PagedAttention memory management, TensorRT-LLM, continuous batching, and cost optimization.