AI & Machine Learning Engineering Notes

· AI & ML · By Zeeshan Ahmad

Comprehensive engineering and research notes covering foundational mathematics, competitive learning, neural gas, PCA, linear models, SVMs, Gaussian processes, and modern LLM / RAG architectures.


#1. Roadmap & Mathematical Foundations

Machine learning involves computer algorithms that improve automatically through experience. The primary goal is for programs to perform complex predictive, classification, and generative tasks by learning directly from data distributions rather than through rigid, hard-coded rules.

#Core Mathematical Disciplines

  • Linear Algebra: Vectors, matrices, transformations, inner products, and coordinate systems for high-dimensional feature spaces.
  • Multivariable Calculus: Derivatives, partial derivatives, Jacobian matrices, Hessian normal forms, and gradients for backpropagation.
  • Probability & Statistics: Covariance matrices, joint and marginal probability distributions, Bayes' Theorem, variance, and standard deviation.

#Essential ML Terminologies

  • Imbalanced Dataset: A dataset where class distributions are heavily skewed. Requires resampling techniques (e.g. SMOTE), class-weighted loss, or precision-recall optimization.
  • Data Partitions:
  • Training Set: Used to fit model parameters.
  • Validation Set: Used for hyperparameter tuning and model selection.
  • Test Set: Completely unseen data used exclusively for unbiased final evaluation.
  • Overfitting vs. Underfitting:
  • Overfitting: High variance; the model memorizes training noise and fails on unseen test distributions.
  • Underfitting: High bias; the model lacks expressiveness to capture fundamental data patterns.
  • Euclidean Distance: Straight-line geometric distance in $\mathbb{R}^n$:
$d(x, y) = \sqrt{\sum_{i=1}^n (x_i - y_i)^2} = \|x - y\|_2$
  • Valid Metric Properties: Any valid distance metric $d(x, y)$ must satisfy:
1. Non-negativity: $d(x, y) \ge 0$ 2. Identity of indiscernibles: $d(x, y) = 0 \iff x = y$ 3. Symmetry: $d(x, y) = d(y, x)$ 4. Triangle inequality: $d(x, z) \le d(x, y) + d(y, z)$
  • Eigenvectors & Eigenvalues: Linear transformation scaling factors:
$A v = \lambda v$
  • Confusion Matrix: Evaluation matrix measuring True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN). Key metrics include:
$\text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}}, \quad \text{Recall} = \frac{\text{TP}}{\text{TP} + \text{FN}}, \quad F_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}$

#2. Competitive Learning & K-Means Clustering

#Clustering Formulation

Clustering groups an unlabeled dataset into $K$ clusters such that objects within the same cluster minimize within-cluster point scatter: $W(I_1, \dots, I_K) = \sum_{k=1}^K \sum_{i \in I_k} \sum_{j \in I_k} \|x_i - x_j\|^2$

#K-Means Clustering Algorithm

K-Means partitions $n$ observations into $K$ distinct, non-overlapping clusters: 1. Initialize: Assign $K$ initial centroid positions $c_1, \dots, c_K$. 2. Assignment Step: Assign each observation $x_i$ to the nearest centroid $c_k$: $I_k = \{ i : \|x_i - c_k\|^2 \le \|x_i - c_j\|^2, \; \forall j \}$ 3. Update Step: Recompute each cluster center as the arithmetic mean: $c_k = \frac{1}{|I_k|} \sum_{i \in I_k} x_i$ 4. Objective (Sum of Squared Errors): $S(I_1, \dots, I_K) = \sum_{k=1}^K \sum_{i \in I_k} \|x_i - c_k\|^2$

#Simple Competitive Learning (Vector Quantization)

Uses fewer codevectors $c_k$ than data points to represent the data distribution, iteratively updating the winning prototype to minimize quantization error:
  • Quantization Error:
$E_{\text{quant}} = \sum_{i=1}^n \min_k \|x_i - c_k\|^2$
  • Codevector Update Rule:
$c^_{\text{new}} = c^ + \lambda (x - c^)$

#3. Self-Organizing Maps (SOM) & Neural Gas

#Neural Gas (NG) Algorithm

Introduced by Martinetz & Schulten, Neural Gas models the topological structure of complex, non-linear data distributions using a flexible network of codevectors without requiring a rigid pre-defined grid.
bash
[Input Point x] ---> Rank Codevectors by Distance: (w_0 closest, w_1, ..., w_n)
                      |
                      v
Update all codevectors based on Rank k_i:
w_i^(new) = w_i + mu * exp(-k_i / lambda) * (x - w_i)
  • Soft Winner-Take-All Update Rule:
$w_i^{\text{new}} = w_i + \mu \cdot \exp\left(-\frac{k_i}{\lambda}\right) (x - w_i)$
  • $\mu$: Learning rate (decays over time).
  • $k_i$: Rank index of codevector $w_i$ based on proximity to $x$ ($k_i = 0$ for nearest).
  • $\lambda$: Smoothing parameter controlling neighborhood influence.
  • Topological Edge Aging: Connections between top-ranking codevectors are aged; edges older than $a_{\text{max}}$ are pruned. Increasing $a_{\text{max}}$ yields a denser connectivity graph.

#Growing Neural Gas (GNG)

Extends Neural Gas by dynamically inserting new codevectors in areas with high accumulated representation error: $w_r = \frac{w_q + w_f}{2}$ Where $w_q$ is the node with maximum accumulated error and $w_f$ is its highest-error neighbor.

#Kohonen Self-Organizing Maps (SOMs)

Fits a low-dimensional grid (1D/2D lattice) to high-dimensional data, preserving topological spatial adjacency: $w_j^{\text{new}} = w_j + \lambda(t) \cdot h_{ij}(t) \cdot (x - w_j)$ Where $h_{ij}(t) = \exp\left(-\frac{\|r_i - r_j\|^2}{2\sigma^2(t)}\right)$ is the neighborhood function.

#4. Dimensionality Reduction: PCA vs. ICA

Dimensionality reduction simplifies high-dimensional feature spaces while preserving essential data patterns and reducing computational overhead.

bash
+--------------------------------------------------------------------------------+
|                        Dimensionality Reduction                                |
+-----------------------------------+--------------------------------------------+
                                    |
          +-------------------------+-------------------------+
          |                                                   |
+-----------------------------------+   +----------------------------------------+
| Principal Component Analysis (PCA)|   | Independent Component Analysis (ICA)   |
+-----------------------------------+   +----------------------------------------+
| * Second-order moments (Variance) |   | * Higher-order statistics (Kurtosis)   |
| * Orthogonal linear projection    |   | * Non-orthogonal independent axes      |
| * Decorrelates features           |   | * Blind source signal separation       |
| * Eigen-decomposition of Cov(X)   |   | * Non-Gaussian source extraction       |
+-----------------------------------+   +----------------------------------------+

#Steps in Principal Component Analysis (PCA)

1. Mean Centering: Subtract the mean vector $\mu$ from the dataset: $X_c = X - \mu$. 2. Covariance Matrix Computation: $\Sigma = \frac{1}{n-1} X_c^T X_c$ 3. Eigen-Decomposition: $\Sigma v_i = \lambda_i v_i$ 4. Projection: Sort eigenvectors by eigenvalues $\lambda_1 \ge \lambda_2 \ge \dots \ge \lambda_d$ and project data: $Y = X_c W$.

#PCA Implementation in Python

python
import numpy as np
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler

# Standardize dataset
X_scaled = StandardScaler().fit_transform(X)

# Retain 95% of cumulative explained variance
pca = PCA(n_components=0.95)
X_pca = pca.fit_transform(X_scaled)

print(f"Original shape: {X.shape}, Reduced shape: {X_pca.shape}")
print(f"Explained variance ratio: {pca.explained_variance_ratio_}")

#Hierarchical Clustering & Dendrograms

  • Agglomerative (Bottom-Up): Successively merges the two closest clusters according to linkage criteria (Single, Complete, Average, Ward's minimum variance).
  • Metric Sensitivity: Changing from $L_2$ (Euclidean) to $L_1$ (Manhattan) or rescaling feature variances alters the resulting tree topology.

#5. Linear Models & Supervised Regression

#Simple & Multiple Linear Regression

Models the linear relationship between independent predictors $X$ and dependent target $y$: $y = w^T X + b + \epsilon$

#Cost Function & Gradient Descent

  • Mean Squared Error (MSE):
$J(w, b) = \frac{1}{2n} \sum_{i=1}^n (f_{w,b}(x^{(i)}) - y^{(i)})^2$
  • Gradient Update:
$w := w - \alpha \frac{\partial J}{\partial w} = w - \frac{\alpha}{n} \sum_{i=1}^n (f(x^{(i)}) - y^{(i)}) x^{(i)}$ $b := b - \alpha \frac{\partial J}{\partial b} = b - \frac{\alpha}{n} \sum_{i=1}^n (f(x^{(i)}) - y^{(i)})$

#6. Support Vector Machines (SVM) & Kernel Methods

#Hard-Margin SVM Formulation

Finds the optimal separating hyperplane that maximizes the geometric margin $\frac{2}{\|w\|}$: $\min_{w, b} \frac{1}{2} \|w\|^2 \quad \text{subject to} \quad y_i (w^T x_i - b) \ge 1, \; \forall i$
  • Support Vectors: The critical training instances lying directly on the margin boundaries where $y_i (w^T x_i - b) = 1$. They represent the active constraints in the optimization problem.
bash
                    Class +1  (y = +1)
                      o     o
                 o      o [SV]  <--- Margin boundary: w^T x - b = +1
               -----------------------
               \                     /
                \  Decision Boundary /   w^T x - b = 0  (Separating Hyperplane)
                 \  Margin = 2/||w|| /
               -----------------------
                 x     [SV]   x <--- Margin boundary: w^T x - b = -1
                     x     x
                    Class -1  (y = -1)

#Common Kernel Functions

When data is not linearly separable in the original input space, kernel functions project inputs into a higher-dimensional reproducing kernel Hilbert space (RKHS):
  • Linear Kernel: $K(x, y) = x^T y$
  • Polynomial Kernel: $K(x, y) = (\gamma x^T y + c)^d$
  • Radial Basis Function (RBF / Gaussian):
$K(x, y) = \exp(-\gamma \|x - y\|^2)$

#7. Gaussian Processes (GP) & Reproducible ML

#Gaussian Process Regression

A Gaussian Process is a non-parametric Bayesian framework defining a prior distribution over functions: $f(x) \sim \mathcal{GP}(m(x), k(x, x'))$
  • Given training observations $(X, y)$, predictions at query points $x^$ yield both a deterministic mean prediction $\mu(x^)$ and an analytic uncertainty variance $\sigma^2(x^)$.

#Pillars of Reproducible Machine Learning

1. Automated Pipelines: Replacing manual experimentation steps with scripted, deterministic execution workflows. 2. Environment & Dependency Pinning: Strict specification of software libraries, GPU drivers, and seed initialization. 3. Open Artifacts & Checkpoints: Publishing model weights, configuration parameters, and dataset splits.

#8. Modern NLP, Vector Embeddings & Production RAG

bash
[Raw Documents / PDFs]
        |
        v
[Recursive Character Text Splitter] (chunk_size=1000, overlap=50)
        |
        v
[Local Embedding Model: nomic-embed-text / text-embedding-3]
        |
        v
[Vector Database: Pinecone / Milvus / Chroma]
        |
        +-----> [User Query] ---> [Cosine Vector Search] ---> [Top-K Context]
                                                                   |
                                                                   v
                                                      [LLM Response Generation]

#Production RAG Implementation with LangChain & Pinecone

python
import os
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import PyPDFDirectoryLoader
from langchain_community.embeddings import OllamaEmbeddings
from pinecone import Pinecone, ServerlessSpec

# 1. Document ingestion and chunking
loader = PyPDFDirectoryLoader("./documents")
raw_docs = loader.load()

text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=50
)
chunks = text_splitter.split_documents(raw_docs)

# 2. Local vector embeddings via Ollama
embeddings = OllamaEmbeddings(model="nomic-embed-text")

# 3. Vector Database upsert
pc = Pinecone(api_key=os.getenv("PINECONE_API_KEY"))
index = pc.Index("engineering-knowledge-base")

batch_size = 100
for i in range(0, len(chunks), batch_size):
    batch = chunks[i : i + batch_size]
    vectors = [
        {
            "id": f"chunk-{i + idx}",
            "values": embeddings.embed_query(doc.page_content),
            "metadata": {"text": doc.page_content, "source": doc.metadata.get("source", "")}
        }
        for idx, doc in enumerate(batch)
    ]
    index.upsert(vectors=vectors)

# 4. Contextual retrieval
query_vector = embeddings.embed_query("What are the mathematical properties of SVM support vectors?")
results = index.query(vector=query_vector, top_k=5, include_metadata=True)

for match in results["matches"]:
    print(f"Similarity Score: {match['score']:.4f}")
    print(f"Content:\n{match['metadata']['text']}\n" + "-" * 50)

#9. Essential Command Reference

#Database & Background Services

bash
# Check PostgreSQL status & interactive terminal
psql -h localhost -U apple -d postgres -c "\l"

# Grant database & schema permissions
psql -h localhost -U apple -d postgres -c "CREATE DATABASE ml_platform;"
psql -h localhost -U apple -d ml_platform -c "GRANT ALL PRIVILEGES ON SCHEMA public TO app_user;"

# Port process management
lsof -ti:5000 | xargs kill -9

#Local LLM Engine (Ollama)

bash
# Manage local model weights
ollama list
ollama run codegemma
ollama run nomic-embed-text

Related articles

Home · All Tools · Blog