{ "cells": [ { "cell_type": "markdown", "id": "d44cffb1", "metadata": {}, "source": [ "# NumPy\n", "\n", "A comprehensive guide to NumPy operations with practical examples.\n", "\n", "References:\n", "- https://numpy.org/doc/\n", "- https://www.w3schools.com/python/numpy/numpy_array_slicing.asp" ] }, { "cell_type": "markdown", "id": "5be9e235", "metadata": {}, "source": [ "## Basics\n", "\n", "Key points from the NumPy docs:\n", "\n", "- Fixed size; resizing creates a new array.\n", "- Same data type for all elements.\n", "- Element-wise ops use compiled C code (fast).\n", "- Vectorization: no explicit loops or indexing.\n", "- Broadcasting: implicit element-wise ops on arrays of different shapes.\n", "- Advanced indexing returns a copy; basic slicing returns a view. `x[(1,2,3),]` ≠ `x[1,2,3]`.\n", "- Python arrays with tuples can be converted to multi-dimensional arrays.\n", "- Specify data type at creation: `np.array([], dtype=complex)`.\n", "- Changing size is expensive; use `np.zeros`, `np.ones` for fixed-length arrays, e.g., `np.ones(3, dtype=np.int32)`.\n", "- Random: `rg = np.random.default_rng(1)`; `rg.random((2,3))`.\n", "- Upcasting: combining arrays of different types yields a more general type.\n" ] }, { "cell_type": "code", "execution_count": 1, "id": "b02b1ba2", "metadata": {}, "outputs": [], "source": [ "# HIDDEN\n", "import numpy as np\n", "import pandas as pd\n", "# Helper function to display matrix\n", "def display(df, indent=1):\n", " s = df.to_string(index=False, header=False)\n", " indented = \"\\n\".join(\" \" * indent + line for line in s.splitlines())\n", " print(indented)\n", " print(\"\")\n", "\n", "def eq(a, b):\n", " assert np.array_equal(a, np.array(b)), f\"\\nExpected:\\n{b}\\nGot:\\n{a}\"\n" ] }, { "cell_type": "markdown", "id": "c0374be0", "metadata": {}, "source": [ "## Matrix operations\n", "\n", "Here is a quick API summary\n", "\n", "```\n", "X.flatten() # Flatten\n", "np.sqrt(X) # Square root all elements\n", "np.sum(X) # Sum all elements\n", "np.sum(X,axis=0) # Row-wise sum\n", "np.sum(X,axis=1) # Column-wise sum\n", "np.amax(X) # Single max value\n", "np.amax(X, axis=0) # Get max in each column\n", "np.amax(X, axis=1) # Get max in each row\n", "np.mean(X) # Mean\n", "np.std(X) # Standard deviation\n", "np.var(X) # Variance\n", "np.trace(X) # Sum of the elements on the diagonal\n", "np.linalg.matrix_rank(X) # Rank of the matrix\n", "np.linalg.det(X) # Determinant of the matrix\n", "```" ] }, { "cell_type": "markdown", "id": "29ad3de5", "metadata": {}, "source": [ "## Slicing\n", "\n", "### 1D slicing\n" ] }, { "cell_type": "markdown", "id": "ca4a2eb1", "metadata": {}, "source": [ "Basic 1D slicing follows `[start:stop:step]` syntax:" ] }, { "cell_type": "code", "execution_count": null, "id": "53fc4780", "metadata": {}, "outputs": [], "source": [ "x = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])\n", "\n", "# Basic slicing\n", "eq(x[2:7], [2, 3, 4, 5, 6]) # Elements 2 through 6\n", "eq(x[:5], [0, 1, 2, 3, 4]) # First 5 elements\n", "eq(x[5:], [5, 6, 7, 8, 9]) # From index 5 to end\n", "eq(x[-3:], [7, 8, 9]) # Last 3 elements\n", "\n", "# With step\n", "eq(x[::2], [0, 2, 4, 6, 8]) # Every other element\n", "eq(x[1::2], [1, 3, 5, 7, 9]) # Every other, starting from 1\n", "eq(x[::-1], [9, 8, 7, 6, 5, 4, 3, 2, 1, 0]) # Reverse\n", "eq(x[8:2:-1], [8, 7, 6, 5, 4, 3]) # Reverse from 8 to 3" ] }, { "cell_type": "markdown", "id": "edf4d2f7", "metadata": {}, "source": [ "#### Boolean indexing\n", "\n", "Select elements based on conditions:" ] }, { "cell_type": "code", "execution_count": null, "id": "24ac4081", "metadata": {}, "outputs": [], "source": [ "x = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])\n", "\n", "# Boolean conditions\n", "eq(x[x > 5], [6, 7, 8, 9])\n", "eq(x[x % 2 == 0], [0, 2, 4, 6, 8]) # Even numbers\n", "eq(x[(x > 2) & (x < 7)], [3, 4, 5, 6]) # AND condition\n", "eq(x[(x < 3) | (x > 7)], [0, 1, 2, 8, 9]) # OR condition\n", "\n", "# Modify elements matching condition\n", "y = x.copy()\n", "y[y > 5] = 0\n", "eq(y, [0, 1, 2, 3, 4, 5, 0, 0, 0, 0])" ] }, { "cell_type": "markdown", "id": "76f901d3", "metadata": {}, "source": [ "## Advanced indexing\n", "\n", "Unlike basic slicing using `:` which returns a \"view\", advanced indexing uses integer arrays, list, or boolean masks.\n" ] }, { "cell_type": "markdown", "id": "666b77a7", "metadata": {}, "source": [ "#### Get specific columes, rows, values from 2D" ] }, { "cell_type": "code", "execution_count": 2, "id": "0290b984", "metadata": {}, "outputs": [], "source": [ "a = np.arange(9).reshape(3,3)\n", "\n", "# Arbitrary element selection\n", "eq(a[[2,0,1], [1,2,0]], [7, 2, 3]) # (2,1)->7, (0,2)->2, (1,0)->3\n", "\n", "# Specific rows\n", "eq(a[[0,2]], [[0, 1, 2],\n", " [6, 7, 8]])\n", "\n", "# Specific columns\n", "eq(a[:, [0,2]], [[0, 2],\n", " [3, 5],\n", " [6, 8]])\n", "\n", "# Boolean mask for elements\n", "mask = a > 4\n", "eq(a[mask], [5, 6, 7, 8]) # flattened selection\n" ] }, { "cell_type": "markdown", "id": "a2a22cf6", "metadata": {}, "source": [ "#### Indexing with steps" ] }, { "cell_type": "code", "execution_count": 3, "id": "5d2914b9", "metadata": {}, "outputs": [], "source": [ "# i:k:l - i is the starting, j the end, k the step.\n", "x = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])\n", "eq(x[1:7:2], [1, 3, 5])\n", "\n" ] }, { "cell_type": "markdown", "id": "72c93d1e", "metadata": {}, "source": [ "### 2D slicing\n", "\n", "Slicing returns a view, rather than a copy.\n", "\n", "```\n", "X = np.arrange(0, 10, 1).reshape((3, 3))\n", "X[1, :] # get second row\n", "X[:, -1] # get last col\n", "X[0:2, :] # get first two rows\n", "X[[0, 2], :] # get first and third rows\n", "X[:, 0:2] # get first two columns\n", "X[:, [0, 2]] # get first and third columns\n", "X[0:2, 0:2] # get submatrix of first two rows/columns\n", "X[X > 5] # get elements greater than 5\n", "```" ] }, { "cell_type": "markdown", "id": "ee8c2837", "metadata": {}, "source": [ "If there are two `:` in each row or column, this implies that `k` value can be now adjusted. In other words, you can selectively choose the internal either across rows or columns, respectively." ] }, { "cell_type": "code", "execution_count": 4, "id": "87b7ee01", "metadata": {}, "outputs": [], "source": [ "arr = np.array([[ 1, 2, 3, 4, 5],\n", " [ 6, 7, 8, 9, 10],\n", " [11, 12, 13, 14, 15]])\n", "\n", "# 2. Slice with steps\n", "\n", "# a) Every 2nd column (step of 2), i,j=all, k=2\n", "eq(arr[:, ::2], [[ 1, 3, 5],\n", " [ 6, 8, 10],\n", " [11, 13, 15]])\n", "\n", "# b) Every 2nd column starting from col=1, i=1, j=all, k=2\n", "eq(arr[:, 1::2], [[ 2, 4],\n", " [ 7, 9],\n", " [12, 14]])\n", "\n", "# c) Every 2nd row (started from i=0)\n", "eq(arr[::2, :], [[ 1, 2, 3, 4, 5],\n", " [11, 12, 13, 14, 15]])\n", "\n", "# d) Every row, but every 3rd column only\n", "eq(arr[:, ::3], [[ 1, 4],\n", " [ 6, 9],\n", " [11, 14]])\n", "\n", "# e) From 2nd row to end, every 2nd column\n", "eq(arr[1:, ::2], [[ 6, 8, 10],\n", " [11, 13, 15]])\n", "\n", "# f) From 1st row to last, every 2nd row and 2nd column\n", "eq(arr[::2, ::2], [[ 1, 3, 5],\n", " [11, 13, 15]])\n", "\n", "# g) Reverse every 2nd column (step -2)\n", "eq(arr[:, ::-2], [[ 5, 3, 1],\n", " [10, 8, 6],\n", " [15, 13, 11]])\n", "\n", "# h) Reverse rows, but only every other one (step -2)\n", "eq(arr[::-2, :], [[11, 12, 13, 14, 15],\n", " [ 1, 2, 3, 4, 5]])\n", "\n", "# i) Take a sub-slice with start, stop, and step\n", "eq(arr[0:3:2, 1:5:2], [[ 2, 4],\n", " [12, 14]])\n", "\n", "# 3. Fancy indexing (list/array of indices)\n", "eq(arr[[0, 1], [0, 2]], [1, 8]) # elements (0,0) and (1,2)\n", "eq(arr[[1, 0], [4, 0]], [10, 1]) # (1,4) and (0,0)\n", "\n", "# 4. Using np.ix_ for 2D fancy indexing\n", "rows = [0, 1]\n", "cols = [1, 3, 4]\n", "eq(arr[np.ix_(rows, cols)],\n", " [[2, 4, 5],\n", " [7, 9, 10]]) # cross-product of row/col indices\n", "\n" ] }, { "cell_type": "markdown", "id": "25ac73a3", "metadata": {}, "source": [ "Reverse rows and columns" ] }, { "cell_type": "code", "execution_count": 5, "id": "1d71b20d", "metadata": {}, "outputs": [ { "ename": "AssertionError", "evalue": "\nExpected:\n[[5, 4, 3, 2, 1], [10, 9, 8, 7, 6], [11, 12, 13, 14, 15]]\nGot:\n[[ 5 4 3 2 1]\n [10 9 8 7 6]\n [15 14 13 12 11]]", "output_type": "error", "traceback": [ "\u001b[0;31m---------------------------------------------------------------------------\u001b[0m", "\u001b[0;31mAssertionError\u001b[0m Traceback (most recent call last)", "Cell \u001b[0;32mIn[5], line 6\u001b[0m\n\u001b[1;32m 1\u001b[0m \u001b[38;5;66;03m# 1. Reverse rows or columns\u001b[39;00m\n\u001b[1;32m 2\u001b[0m eq(arr[::\u001b[38;5;241m-\u001b[39m\u001b[38;5;241m1\u001b[39m, :], [[\u001b[38;5;241m11\u001b[39m, \u001b[38;5;241m12\u001b[39m, \u001b[38;5;241m13\u001b[39m, \u001b[38;5;241m14\u001b[39m, \u001b[38;5;241m15\u001b[39m],\n\u001b[1;32m 3\u001b[0m [\u001b[38;5;241m6\u001b[39m, \u001b[38;5;241m7\u001b[39m, \u001b[38;5;241m8\u001b[39m, \u001b[38;5;241m9\u001b[39m, \u001b[38;5;241m10\u001b[39m],\n\u001b[1;32m 4\u001b[0m [\u001b[38;5;241m1\u001b[39m, \u001b[38;5;241m2\u001b[39m, \u001b[38;5;241m3\u001b[39m, \u001b[38;5;241m4\u001b[39m, \u001b[38;5;241m5\u001b[39m]]) \u001b[38;5;66;03m# reverse rows\u001b[39;00m\n\u001b[0;32m----> 6\u001b[0m \u001b[43meq\u001b[49m\u001b[43m(\u001b[49m\u001b[43marr\u001b[49m\u001b[43m[\u001b[49m\u001b[43m:\u001b[49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[43m:\u001b[49m\u001b[43m:\u001b[49m\u001b[38;5;241;43m-\u001b[39;49m\u001b[38;5;241;43m1\u001b[39;49m\u001b[43m]\u001b[49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[43m[\u001b[49m\u001b[43m[\u001b[49m\u001b[38;5;241;43m5\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m4\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m3\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m2\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m1\u001b[39;49m\u001b[43m]\u001b[49m\u001b[43m,\u001b[49m\n\u001b[1;32m 7\u001b[0m \u001b[43m \u001b[49m\u001b[43m[\u001b[49m\u001b[38;5;241;43m10\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m9\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m8\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m7\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m6\u001b[39;49m\u001b[43m]\u001b[49m\u001b[43m,\u001b[49m\n\u001b[1;32m 8\u001b[0m \u001b[43m \u001b[49m\u001b[43m[\u001b[49m\u001b[38;5;241;43m11\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m12\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m13\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m14\u001b[39;49m\u001b[43m,\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m15\u001b[39;49m\u001b[43m]\u001b[49m\u001b[43m]\u001b[49m\u001b[43m)\u001b[49m \u001b[38;5;66;03m# reverse columns\u001b[39;00m\n", "Cell \u001b[0;32mIn[1], line 12\u001b[0m, in \u001b[0;36meq\u001b[0;34m(a, b)\u001b[0m\n\u001b[1;32m 11\u001b[0m \u001b[38;5;28;01mdef\u001b[39;00m\u001b[38;5;250m \u001b[39m\u001b[38;5;21meq\u001b[39m(a, b):\n\u001b[0;32m---> 12\u001b[0m \u001b[38;5;28;01massert\u001b[39;00m np\u001b[38;5;241m.\u001b[39marray_equal(a, np\u001b[38;5;241m.\u001b[39marray(b)), \u001b[38;5;124mf\u001b[39m\u001b[38;5;124m\"\u001b[39m\u001b[38;5;130;01m\\n\u001b[39;00m\u001b[38;5;124mExpected:\u001b[39m\u001b[38;5;130;01m\\n\u001b[39;00m\u001b[38;5;132;01m{\u001b[39;00mb\u001b[38;5;132;01m}\u001b[39;00m\u001b[38;5;130;01m\\n\u001b[39;00m\u001b[38;5;124mGot:\u001b[39m\u001b[38;5;130;01m\\n\u001b[39;00m\u001b[38;5;132;01m{\u001b[39;00ma\u001b[38;5;132;01m}\u001b[39;00m\u001b[38;5;124m\"\u001b[39m\n", "\u001b[0;31mAssertionError\u001b[0m: \nExpected:\n[[5, 4, 3, 2, 1], [10, 9, 8, 7, 6], [11, 12, 13, 14, 15]]\nGot:\n[[ 5 4 3 2 1]\n [10 9 8 7 6]\n [15 14 13 12 11]]" ] } ], "source": [ "# 1. Reverse rows or columns\n", "eq(arr[::-1, :], [[11, 12, 13, 14, 15],\n", " [6, 7, 8, 9, 10],\n", " [1, 2, 3, 4, 5]]) # reverse rows\n", "\n", "eq(arr[:, ::-1], [[5, 4, 3, 2, 1],\n", " [10, 9, 8, 7, 6],\n", " [11, 12, 13, 14, 15]]) # reverse columns\n" ] }, { "cell_type": "markdown", "id": "1f0c83d3", "metadata": {}, "source": [ "### Slicing with intervals" ] }, { "cell_type": "code", "execution_count": null, "id": "e4ae45ed", "metadata": {}, "outputs": [], "source": [ "arr = np.array([[1, 2, 3, 4, 5],\n", " [6, 7, 8, 9, 10]])\n", "\n", "eq(arr[0:1, 1:4], np.array([[2, 3, 4]]))\n", "eq(arr[:1, 1:4], np.array([[2, 3, 4]]))\n", "eq(arr[0:2, 2], np.array([3, 8]))\n", "eq(arr[0:2, 1:4], np.array([[2, 3, 4],\n", " [7, 8, 9]]))\n" ] }, { "cell_type": "markdown", "id": "a3aabfeb", "metadata": {}, "source": [ "## Add new dimension\n", "\n", "``none`` is used to insert a new axis or dimension." ] }, { "cell_type": "code", "execution_count": null, "id": "98552b34", "metadata": {}, "outputs": [], "source": [ "arr = np.arange(10) \n", "assert arr.shape == (10,)\n", "# Add two new axes using [:, None, None]\n", "reshaped = arr[:, None, None]\n", "assert reshaped.shape == (10, 1, 1)" ] }, { "cell_type": "code", "execution_count": null, "id": "4d50565d", "metadata": {}, "outputs": [], "source": [ "arr2 = arr.reshape(2, 5)\n", "assert arr2.shape == (2, 5)\n", "# Add two new axes after the first axis (\"row\")\n", "assert arr2[:, None, None].shape == (2, 1, 1, 5)\n", "assert arr2[:, :, None, None].shape == (2, 5, 1, 1)" ] }, { "cell_type": "markdown", "id": "90ccef3b", "metadata": {}, "source": [ "## Create and copy tensor" ] }, { "cell_type": "code", "execution_count": null, "id": "de8d10f0", "metadata": {}, "outputs": [], "source": [ "eq(np.matrix(np.arange(12).reshape((3, 4))), np.arange(12).reshape(3, 4))\n", "eq(np.zeros((5,), dtype=int), np.array([0, 0, 0, 0, 0], dtype=int))\n", "eq(np.zeros((2, 1)), np.zeros((2, 1)))\n", "\n", "# Reshape\n", "X = np.arange(6).reshape((2, 3))\n", "eq(X, np.array([[0, 1, 2],\n", " [3, 4, 5]]))\n", "\n", "# Copy exactly\n", "eq(np.copy(X), X)\n", "\n", "# Copy shape\n", "eq(np.ones_like(X), np.array([[1, 1, 1],\n", " [1, 1, 1]]))\n", "eq(np.zeros_like(X), np.array([[0, 0, 0],\n", " [0, 0, 0]]))\n", "\n", "# Full\n", "eq(np.full((2, 2), 10), np.array([[10, 10],\n", " [10, 10]]))\n", "eq(np.full((2, 2), np.inf), np.array([[np.inf, np.inf],\n", " [np.inf, np.inf]]))\n", "eq(np.full((2, 2), [1, 2]), np.array([[1, 2],\n", " [1, 2]]))" ] }, { "cell_type": "markdown", "id": "bfdeb541", "metadata": {}, "source": [ "## Broadcast" ] }, { "cell_type": "code", "execution_count": null, "id": "a77274b9", "metadata": {}, "outputs": [], "source": [ "a = np.array([1,2,3])\n", "assert a.shape == (3,)\n", "b = np.array([\n", " [10],\n", " [20],\n", " [30]])\n", "assert b.shape == (3,1)\n", "\n", "# In a, 1, 2, 3 duplicated across new rows\n", "# In b, 10, 20, 30 duplicated acorss new columns\n", "# And then those are added\n", "expected = np.array([[11,12,13],\n", " [21,22,23],\n", " [31,32,33]])\n", "\n", "assert np.array_equal(a + b, expected)" ] }, { "cell_type": "markdown", "id": "bab2db64", "metadata": {}, "source": [ "## Stacking\n", "\n", "- Axis 0 - rows\n", "- Axis 1 - columns\n", "- Axis 2 - depth\n", "- Axis 3 - so on.." ] }, { "cell_type": "code", "execution_count": null, "id": "da7ba3c2", "metadata": {}, "outputs": [], "source": [ "# Base arrays\n", "a = np.array([1, 2, 3])\n", "b = np.array([4, 5, 6])\n", "\n", "# Stack across rows (Method 1)\n", "eq(np.stack([a, b], axis=0),\n", " np.array([[1, 2, 3],\n", " [4, 5, 6]]))\n", "\n", "# Stack across rows (Method 2)\n", "eq(np.vstack([a, b]),\n", " np.array([[1, 2, 3],\n", " [4, 5, 6]]))\n", "\n", "# Stack across columns (like rotate 90 and row-sum)\n", "eq(np.stack([a, b], axis=1),\n", " np.array([[1, 4],\n", " [2, 5],\n", " [3, 6]]))\n", "\n", "# Concatenate along columns\n", "eq(np.hstack([a, b]),\n", " np.array([1, 2, 3, 4, 5, 6]))\n", "\n", "# Stack along depth / third axis\n", "c = np.array([7, 8, 9])\n", "eq(np.dstack([a, b, c]),\n", " np.array([[[1, 4, 7],\n", " [2, 5, 8],\n", " [3, 6, 9]]]))" ] }, { "cell_type": "markdown", "id": "7b9d8f3f", "metadata": {}, "source": [ "Just to note that `np.vstack` is a shorthand for vertical stacking `like np.concatenate(..., axis=0)`. `np.stack` lets you choose any axis so it's more general." ] }, { "cell_type": "markdown", "id": "91095d7b", "metadata": {}, "source": [ "## Performance\n", "\n", "- Vectoization - use array ops to loops\n", "- use ``where`` for conditional element selection ``np.where(X > 5, 1, 0) # Replace with 1 if >5 else 0``" ] }, { "cell_type": "code", "execution_count": null, "id": "1aaa0c10", "metadata": {}, "outputs": [], "source": [ "# Example array with NaN and Inf\n", "arr = np.array([1.0, 2.0, np.nan, np.inf, -np.inf, 3.0])\n", "\n", "# Count NaNs\n", "assert np.isnan(arr).sum() == 1 # only one np.nan\n", "\n", "# Count Infs\n", "assert np.isinf(arr).sum() == 2 # +inf and -inf\n", "\n", "# Mean ignoring NaNs\n", "arr2 = np.array([1.0, 2.0, np.nan, 3.0])\n", "assert np.nanmean(arr2) == 2.0 # (1+2+3)/3\n", "\n", "# Replace NaN/Inf with finite values\n", "cleaned = np.nan_to_num(arr, nan=0.0, posinf=999.0, neginf=-999.0)\n", "expected = np.array([1.0, 2.0, 0.0, 999.0, -999.0, 3.0])\n", "assert np.array_equal(cleaned, expected)" ] }, { "cell_type": "code", "execution_count": null, "id": "2fe3b4ee", "metadata": {}, "outputs": [], "source": [ "np.set_printoptions(\n", " precision=3, # Set decimal places\n", " suppress=True, # Avoid scientific notations\n", " threshold=100, # Max number of elements to be printed\n", " linewidth=80,\n", " edgeitems=2 # Show two values per edge when truncated\n", ")" ] }, { "cell_type": "markdown", "id": "3518f55f", "metadata": {}, "source": [ "## Other useful stuff\n", "\n", "### Print nicely\n", "\n" ] }, { "cell_type": "markdown", "id": "a7619756", "metadata": {}, "source": [ "## Linear algebra operations" ] }, { "cell_type": "code", "execution_count": null, "id": "f26befb7", "metadata": {}, "outputs": [], "source": [ "A = np.array([[1, 2], [3, 4]])\n", "B = np.array([[5, 6], [7, 8]])\n", "\n", "# Matrix multiplication (dot product)\n", "eq(A @ B, [[19, 22], [43, 50]])\n", "eq(np.dot(A, B), [[19, 22], [43, 50]])\n", "\n", "# Element-wise multiplication\n", "eq(A * B, [[5, 12], [21, 32]])\n", "\n", "# Transpose\n", "eq(A.T, [[1, 3], [2, 4]])\n", "\n", "# Inverse\n", "A_inv = np.linalg.inv(A)\n", "identity = A @ A_inv\n", "assert np.allclose(identity, np.eye(2)) # Should be identity matrix\n", "\n", "# Eigenvalues and eigenvectors\n", "eigenvalues, eigenvectors = np.linalg.eig(A)\n", "print(f\"Eigenvalues: {eigenvalues}\")\n", "print(f\"Eigenvectors:\\n{eigenvectors}\")\n", "\n", "# Solve linear system Ax = b\n", "b = np.array([1, 2])\n", "x = np.linalg.solve(A, b)\n", "assert np.allclose(A @ x, b) # Verify solution" ] }, { "cell_type": "markdown", "id": "7c8208ff", "metadata": {}, "source": [ "## Array manipulation" ] }, { "cell_type": "code", "execution_count": null, "id": "65be6c6c", "metadata": {}, "outputs": [], "source": [ "x = np.arange(12)\n", "\n", "# Reshape\n", "eq(x.reshape(3, 4), [[0, 1, 2, 3],\n", " [4, 5, 6, 7],\n", " [8, 9, 10, 11]])\n", "\n", "eq(x.reshape(2, 6), [[0, 1, 2, 3, 4, 5],\n", " [6, 7, 8, 9, 10, 11]])\n", "\n", "# Reshape with -1 (auto-calculate dimension)\n", "eq(x.reshape(3, -1), [[0, 1, 2, 3],\n", " [4, 5, 6, 7],\n", " [8, 9, 10, 11]])\n", "\n", "# Flatten\n", "A = np.array([[1, 2, 3], [4, 5, 6]])\n", "eq(A.flatten(), [1, 2, 3, 4, 5, 6])\n", "eq(A.ravel(), [1, 2, 3, 4, 5, 6]) # Similar but returns view when possible\n", "\n", "# Squeeze - remove axes of length 1\n", "B = np.array([[[1], [2], [3]]])\n", "assert B.shape == (1, 3, 1)\n", "C = np.squeeze(B)\n", "assert C.shape == (3,)\n", "eq(C, [1, 2, 3])\n", "\n", "# Expand dimensions\n", "eq(np.expand_dims(C, axis=0), [[1, 2, 3]])\n", "eq(np.expand_dims(C, axis=1), [[1], [2], [3]])" ] }, { "cell_type": "markdown", "id": "7f7176e6", "metadata": {}, "source": [ "## Statistical operations" ] }, { "cell_type": "code", "execution_count": null, "id": "ea402e50", "metadata": {}, "outputs": [], "source": [ "data = np.array([[1, 2, 3],\n", " [4, 5, 6],\n", " [7, 8, 9]])\n", "\n", "# Basic statistics\n", "print(f\"Mean: {np.mean(data)}\") # 5.0\n", "print(f\"Median: {np.median(data)}\") # 5.0\n", "print(f\"Std: {np.std(data)}\") # ~2.58\n", "print(f\"Var: {np.var(data)}\") # ~6.67\n", "\n", "# Axis-wise operations\n", "print(f\"Column means: {np.mean(data, axis=0)}\") # [4, 5, 6]\n", "print(f\"Row means: {np.mean(data, axis=1)}\") # [2, 5, 8]\n", "\n", "# Min/Max\n", "print(f\"Min: {np.min(data)}\") # 1\n", "print(f\"Max: {np.max(data)}\") # 9\n", "print(f\"Argmin: {np.argmin(data)}\") # 0 (index of min)\n", "print(f\"Argmax: {np.argmax(data)}\") # 8 (index of max)\n", "\n", "# Percentiles\n", "print(f\"25th percentile: {np.percentile(data, 25)}\") # 2.5\n", "print(f\"75th percentile: {np.percentile(data, 75)}\") # 7.5\n", "\n", "# Cumulative operations\n", "print(f\"Cumsum: {np.cumsum([1, 2, 3, 4])}\") # [1, 3, 6, 10]\n", "print(f\"Cumprod: {np.cumprod([1, 2, 3, 4])}\") # [1, 2, 6, 24]" ] }, { "cell_type": "markdown", "id": "a209b6e7", "metadata": {}, "source": [ "## Universal functions (ufuncs)\n", "\n", "Element-wise operations that are highly optimized:" ] }, { "cell_type": "code", "execution_count": null, "id": "ba58dbfb", "metadata": {}, "outputs": [], "source": [ "x = np.array([1, 4, 9, 16, 25])\n", "\n", "# Mathematical functions\n", "eq(np.sqrt(x), [1.0, 2.0, 3.0, 4.0, 5.0])\n", "eq(np.square(x), [1, 16, 81, 256, 625])\n", "eq(np.exp([0, 1, 2]), [1.0, np.e, np.e**2])\n", "eq(np.log([1, np.e, np.e**2]), [0.0, 1.0, 2.0])\n", "\n", "# Trigonometric\n", "angles = np.array([0, np.pi/2, np.pi])\n", "assert np.allclose(np.sin(angles), [0, 1, 0])\n", "assert np.allclose(np.cos(angles), [1, 0, -1])\n", "\n", "# Rounding\n", "y = np.array([1.2, 2.5, 3.7, 4.9])\n", "eq(np.round(y), [1.0, 2.0, 4.0, 5.0])\n", "eq(np.floor(y), [1.0, 2.0, 3.0, 4.0])\n", "eq(np.ceil(y), [2.0, 3.0, 4.0, 5.0])\n", "\n", "# Absolute value\n", "z = np.array([-1, -2, 3, -4])\n", "eq(np.abs(z), [1, 2, 3, 4])\n", "\n", "# Clip values\n", "data = np.array([1, 5, 10, 15, 20])\n", "eq(np.clip(data, 5, 15), [5, 5, 10, 15, 15]) # Limit between 5 and 15" ] }, { "cell_type": "markdown", "id": "962d85d0", "metadata": {}, "source": [ "## Sorting and searching" ] }, { "cell_type": "code", "execution_count": null, "id": "db99b110", "metadata": {}, "outputs": [], "source": [ "arr = np.array([3, 1, 4, 1, 5, 9, 2, 6])\n", "\n", "# Sort (returns new sorted array)\n", "eq(np.sort(arr), [1, 1, 2, 3, 4, 5, 6, 9])\n", "\n", "# Argsort (returns indices that would sort the array)\n", "indices = np.argsort(arr)\n", "eq(indices, [1, 3, 6, 0, 2, 4, 7, 5])\n", "eq(arr[indices], [1, 1, 2, 3, 4, 5, 6, 9]) # Verify\n", "\n", "# Sort 2D array\n", "matrix = np.array([[3, 1, 4],\n", " [1, 5, 9],\n", " [2, 6, 5]])\n", "\n", "eq(np.sort(matrix, axis=0), # Sort each column\n", " [[1, 1, 4],\n", " [2, 5, 5],\n", " [3, 6, 9]])\n", "\n", "eq(np.sort(matrix, axis=1), # Sort each row\n", " [[1, 3, 4],\n", " [1, 5, 9],\n", " [2, 5, 6]])\n", "\n", "# Unique values\n", "data = np.array([1, 2, 2, 3, 3, 3, 4])\n", "eq(np.unique(data), [1, 2, 3, 4])\n", "\n", "# Unique with counts\n", "unique, counts = np.unique(data, return_counts=True)\n", "eq(unique, [1, 2, 3, 4])\n", "eq(counts, [1, 2, 3, 1])\n", "\n", "# Where (find indices matching condition)\n", "x = np.array([1, 2, 3, 4, 5])\n", "indices = np.where(x > 2)\n", "eq(indices[0], [2, 3, 4]) # Indices where x > 2" ] }, { "cell_type": "markdown", "id": "7306c55d", "metadata": {}, "source": [ "### Random number geneator" ] }, { "cell_type": "code", "execution_count": null, "id": "ff1e5251", "metadata": {}, "outputs": [], "source": [ "\n", "# Uniform [0,1)\n", "a = np.random.rand(3, 2)\n", "assert a.shape == (3, 2)\n", "assert np.all((a >= 0) & (a < 1))\n", "\n", "# Standard normal (mean ≈ 0, std ≈ 1, but here just shape check)\n", "b = np.random.randn(3, 2)\n", "assert b.shape == (3, 2)\n", "# Values can be any real number, so no bound check\n", "\n", "# Random integers between 0 and 9\n", "c = np.random.randint(0, 10, (2, 3))\n", "assert c.shape == (2, 3)\n", "assert np.all((c >= 0) & (c < 10))\n", "\n", "# Sampling with replacement\n", "d = np.random.choice([1, 2, 3], size=5, replace=True)\n", "assert d.shape == (5,)\n", "assert np.all(np.isin(d, [1, 2, 3]))\n" ] } ], "metadata": { "kernelspec": { "display_name": "ophus-env", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.13.7" } }, "nbformat": 4, "nbformat_minor": 5 }