EDA & Descriptive Statistics
Mean, Median, Variance, Skewness, Kurtosis, IQR outlier filtering, and probability distributions (Gaussian, Poisson, Exponential).
Hypothesis Testing & P-Values
Null vs Alternative Hypothesis (H0/H1), Type I / Type II errors, Student's t-test, ANOVA, and Chi-Square contingency tests.
Feature Engineering & Scaling
StandardScaler, MinMaxScaler, Box-Cox log transformations, One-Hot / Target encoding, and KNN / MICE missing value imputation.
Linear & Polynomial Regression (OLS)
Closed-form Ordinary Least Squares normal equation derivation, R², Adjusted R², RMSE, and Polynomial feature expansion.
Regularization (Ridge/Lasso) & Bias-Variance
L2 Ridge penalty, L1 Lasso feature selection sparsity, ElasticNet, and navigating the Bias-Variance tradeoff curve.
Logistic Regression & ROC-AUC Metrics
Sigmoid odds ratios, Maximum Likelihood Estimation, ROC-AUC curves, Precision-Recall tradeoffs, and F1-score evaluation.
Decision Trees & Random Forests
Shannon Entropy, Gini Impurity splits, Bootstrap Aggregation (Bagging), feature sub-sampling, and Out-of-Bag (OOB) error.
Gradient Boosting & XGBoost Math
Sequential pseudo-residual tree fitting, 2nd-order Taylor expansions, shrinkage learning rates, and LightGBM histogram binning.
Unsupervised Clustering: K-Means & DBSCAN
Lloyd's EM centroid optimization, Within-Cluster Sum of Squares (WCSS), Silhouette scores, and DBSCAN density clustering.
Dimensionality Reduction: PCA & SVD
Covariance matrix eigendecomposition, Singular Value Decomposition (SVD), explained variance ratio scree plots, and t-SNE / UMAP manifolds.
Time Series Forecasting (ARIMA & SARIMAX)
Augmented Dickey-Fuller stationarity tests, Autocorrelation (ACF) / PACF plots, differencing, and seasonal SARIMAX modeling.
MLOps Pipelines & Drift Monitoring
Scikit-Learn ColumnTransformer pipelines, MLflow experiment tracking, Data Drift (KS-Test), FastAPI inference endpoints, and Docker containers.