Python Data Science Library Mastery • Training 01
Article-Training • Core Numerical Computing

NumPy

Think in Arrays: The Numerical Engine Behind Python Data Science

Learn the mental model that makes NumPy useful: arrays, shape, indexing, vectorization, broadcasting, axes, random generation and a real analytical workflow.

Business Question → Array → Shape → Vectorize → Broadcast → Aggregate → Decision
12
18
21
30
15
22
28
35
↓ × 1.05
One rule → every value
8learning modules
24interactive practices
5rapid review questions
50%certificate threshold
Learning target
Build enough NumPy fluency to read, write and explain common numerical array workflows.
Practice progress0 / 24

Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.

MODULE 01
🧠

What NumPy Is — and Why Arrays Change the Game

NumPy is the foundation of numerical computing in Python. Its core object, the ndarray, lets you work with large blocks of numeric data using concise array operations instead of writing a Python loop for every calculation.

👁️
See it this way

Think of a Python list as a row of labeled boxes. Think of a NumPy array as a numeric grid designed to be processed as one object.

Core ideas

  • Fast N-dimensional arrays (ndarray)
  • Vectorized mathematical operations
  • Shape, dtype and axis awareness
  • Foundation for Pandas, SciPy, scikit-learn and much of the scientific Python ecosystem
Try this
import numpy as np

sales = np.array([120, 135, 128, 150])
print(sales.mean())      # 133.25
print(sales * 1.05)      # 5% scenario
✅

Use NumPy when the work is primarily numerical, matrix-like, simulation-oriented, or performance-sensitive. Use Pandas when labels, mixed data types, joins, dates, and table semantics are the main problem.

Practice the decision, not just the syntax

Practice 1
Which object is NumPy built around?
Practice 2
When is NumPy usually the better first tool?
Practice 3
What does vectorization mean here?
MODULE 02
🧱

Create Arrays and Read Their Shape

Before calculating anything, learn to read the array. Four properties tell you a lot: shape, ndim, size and dtype.

👁️
See it this way

Shape answers “how many rows and columns?”; ndim answers “how many axes?”; size answers “how many values?”; dtype answers “what numeric representation?”

Core ideas

  • np.array() from existing values
  • np.zeros() / np.ones() for initialized arrays
  • np.arange() for sequences
  • reshape() to reorganize without changing the number of elements
Try this
a = np.arange(12).reshape(3, 4)
print(a)
print(a.shape)   # (3, 4)
print(a.ndim)    # 2
print(a.size)    # 12
print(a.dtype)
✅

A reshape is valid only when the new dimensions multiply to the same number of elements. Twelve values can become 3×4, 2×6, 1×12, etc.

Practice the decision, not just the syntax

Practice 4
For an array with shape (3, 4), how many elements are there?
Practice 5
Which property reports the number of dimensions?
Practice 6
Which reshape is valid for 12 values?
MODULE 03
🎯

Index, Slice and Filter Without Fear

Indexing selects positions; slicing selects ranges; Boolean masks select values that satisfy a condition. These three skills are the gateway to practical NumPy.

👁️
See it this way

Position → Range → Condition. First point to one value, then a region, then let a rule choose the region for you.

Core ideas

  • arr[0] selects one element
  • arr[1:4] selects a slice
  • matrix[row, column] selects from 2-D arrays
  • arr[arr > threshold] applies a Boolean mask
Try this
response = np.array([18, 42, 27, 61, 35, 75])
print(response[2])
print(response[1:4])
print(response[response > 40])  # only slower cases
✅

Watch out for views: many basic slices refer to the same underlying data. Editing a slice can therefore affect the original array. Use .copy() when you need an independent copy.

Practice the decision, not just the syntax

Practice 7
What does arr[arr > 10] use?
Practice 8
What can happen when you modify a basic slice?
Practice 9
What should you use when you need an independent sliced array?
MODULE 04
⚡

Vectorization: Replace Row-by-Row Thinking

Vectorization means expressing a calculation as an operation on an entire array. NumPy executes many of these operations in optimized compiled code.

👁️
See it this way

Instead of “for each number, do X,” think “apply X to the whole array.” That mental shift is one of the most important NumPy skills.

Core ideas

  • Arithmetic operators act element-by-element
  • Universal functions (ufuncs) include np.sqrt, np.exp, np.maximum and many more
  • Vectorized expressions are usually clearer and often faster than explicit Python loops
  • Avoid np.vectorize as a performance assumption; it is mainly a convenience wrapper
Try this
minutes = np.array([30, 45, 52, 28])
target = 40
variance = minutes - target
status = np.where(minutes <= target, 'On target', 'Late')
print(variance)
print(status)
✅

Vectorization shines when the same numerical rule applies across many values. A loop is still valid when every item requires complex branching, external I/O, or object-level logic.

Practice the decision, not just the syntax

Practice 10
What does minutes - target do when target is a scalar?
Practice 11
Which is a NumPy universal function?
Practice 12
Which statement is safest about np.vectorize?
MODULE 05
📡

Broadcasting: Small Shapes, Big Calculations

Broadcasting lets NumPy combine arrays with compatible shapes without manually copying the smaller array. It is powerful because it expresses rules by dimension.

👁️
See it this way

A scalar can stretch across every cell. A 1-D row can stretch across rows of a 2-D matrix. Compatibility is checked from the trailing dimensions.

Core ideas

  • Dimensions are compatible when they are equal or one of them is 1
  • Broadcasting avoids unnecessary explicit copies in many cases
  • Shape errors are often broadcasting errors in disguise
  • Print .shape before guessing
Try this
sales = np.array([[100, 120, 90],
                  [80, 110, 105]])
factors = np.array([1.00, 1.05, 0.98])
adjusted = sales * factors
print(adjusted)
✅

When two shapes will not broadcast, reshape deliberately. For example, a per-row factor may need factors[:, None] so NumPy sees a column instead of a flat 1-D vector.

Practice the decision, not just the syntax

Practice 13
When are two dimensions broadcasting-compatible?
Practice 14
What is the first diagnostic when broadcasting surprises you?
Practice 15
What does x[:, None] commonly do to a 1-D array?
MODULE 06
📊

Aggregate by Axis — Mean, Sum, Min, Max

Aggregations compress many values into summaries. The key is understanding axis: which dimension disappears after the calculation?

👁️
See it this way

For a 2-D table: axis=0 aggregates down the rows and returns one result per column; axis=1 aggregates across columns and returns one result per row.

Core ideas

  • mean, sum, min, max, std and percentile summarize arrays
  • axis controls the direction of reduction
  • keepdims=True can preserve dimensions for later broadcasting
  • NaN-aware variants such as np.nanmean are useful when missing numeric values are represented by NaN
Try this
daily = np.array([[22, 30, 35],
                  [18, 26, 40],
                  [25, 28, 32]])
print(daily.mean())
print(daily.mean(axis=0))  # per column
print(daily.mean(axis=1))  # per row
✅

If an axis result surprises you, compare the input shape and output shape. The reduced axis is the one removed unless keepdims=True.

Practice the decision, not just the syntax

Practice 16
In a 2-D array, what does mean(axis=0) typically return?
Practice 17
Why use keepdims=True?
Practice 18
Which function is designed to ignore NaN values when computing a mean?
MODULE 07
🎲

Random Numbers and Reproducible Experiments

Random numbers support simulation, sampling, testing and synthetic data. Modern NumPy examples use a Generator created with np.random.default_rng().

👁️
See it this way

Random does not have to mean unrepeatable. A fixed seed creates a reproducible sequence — critical for debugging and teaching.

Core ideas

  • rng = np.random.default_rng(seed)
  • rng.random() for values in [0, 1)
  • rng.integers() for random integers
  • rng.normal() for normal-distribution samples
Try this
rng = np.random.default_rng(42)
simulated_minutes = rng.normal(loc=35, scale=8, size=1000)
print(simulated_minutes.mean())
print(np.percentile(simulated_minutes, [50, 90, 95]))
✅

A seed gives repeatability, not “better randomness.” In analysis, record the seed when reproducibility matters.

Practice the decision, not just the syntax

Practice 19
Which is the modern pattern for creating a random generator?
Practice 20
What does a fixed seed primarily provide?
Practice 21
Which method generates normal-distribution samples from a Generator?
MODULE 08
🏗️

From Array Skills to a Real Data Workflow

The real skill is not memorizing functions. It is recognizing when a problem can be translated into arrays, shapes, masks, vectorized rules, broadcasting and aggregation.

👁️
See it this way

Business question → Numeric array → Clean shape → Vectorized rule → Aggregate → Interpret → Hand off to Pandas / visualization / ML when labels and richer context are needed.

Core ideas

  • Use NumPy to build the numerical engine
  • Use Pandas to add labeled-table workflows
  • Use Matplotlib / Plotly to communicate results
  • Use scikit-learn when the next step is predictive modeling
Try this
# Mini case: service response minutes by team and day
response = np.array([[32, 35, 40, 28],
                     [45, 50, 39, 42],
                     [25, 27, 31, 29]])
target = np.array([35, 40, 30])[:, None]

team_avg = response.mean(axis=1)
within_target = team_avg <= target.ravel()
print(team_avg)
print(within_target)
✅

Market skill: explain the transformation, not just the syntax. Employers value analysts who can connect array operations to measurable decisions, quality checks and downstream models.

Practice the decision, not just the syntax

Practice 22
What is the strongest workflow mindset?
Practice 23
When should Pandas usually take over from NumPy?
Practice 24
What makes a NumPy skill market-relevant?
5-Question Knowledge Check

Can you explain NumPy without looking at syntax?

Open each item only after answering it in your own words.

1. What is the difference between shape and size?

Shape describes dimensions; size is the total number of elements.

2. Why can a slice unexpectedly change the source array?

Because basic slicing commonly creates a view that shares the same underlying data.

3. What problem does broadcasting solve?

It lets compatible shapes participate in the same operation without manually duplicating the smaller array.

4. How do you debug axis confusion?

Write down the input shape, identify the dimension being reduced, and compare the output shape.

5. Why use default_rng(42) in a demo?

A fixed seed makes the pseudo-random results reproducible for teaching, testing and debugging.

Decision Guide

NumPy, Python List or Pandas?

NeedPython ListNumPyPandas
Simple general-purpose container✅○○
Dense numerical computation○✅✅/○
Broadcasting / matrix-style math○✅✅/○
Named columns, joins, dates, mixed types○○✅
Foundation for scientific / ML stack○✅✅
Official Sources & Further Learning

Grounded in current NumPy documentation

The technical concepts in this training follow the NumPy 2.x documentation.

Market Skills

What you should be able to say after this training

“I can convert numerical business rules into NumPy arrays, validate shapes and dtypes, filter with Boolean masks, vectorize calculations, use broadcasting deliberately, summarize by axis, and create reproducible random simulations.”

Certificate of Participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%