Learn the mental model that makes NumPy useful: arrays, shape, indexing, vectorization, broadcasting, axes, random generation and a real analytical workflow.
Complete 12 of 24 practices (50%) and enter your name to unlock the Certificate of Participation.
NumPy is the foundation of numerical computing in Python. Its core object, the ndarray, lets you work with large blocks of numeric data using concise array operations instead of writing a Python loop for every calculation.
Think of a Python list as a row of labeled boxes. Think of a NumPy array as a numeric grid designed to be processed as one object.
import numpy as np
sales = np.array([120, 135, 128, 150])
print(sales.mean()) # 133.25
print(sales * 1.05) # 5% scenarioUse NumPy when the work is primarily numerical, matrix-like, simulation-oriented, or performance-sensitive. Use Pandas when labels, mixed data types, joins, dates, and table semantics are the main problem.
Before calculating anything, learn to read the array. Four properties tell you a lot: shape, ndim, size and dtype.
Shape answers “how many rows and columns?”; ndim answers “how many axes?”; size answers “how many values?”; dtype answers “what numeric representation?”
a = np.arange(12).reshape(3, 4)
print(a)
print(a.shape) # (3, 4)
print(a.ndim) # 2
print(a.size) # 12
print(a.dtype)A reshape is valid only when the new dimensions multiply to the same number of elements. Twelve values can become 3×4, 2×6, 1×12, etc.
Indexing selects positions; slicing selects ranges; Boolean masks select values that satisfy a condition. These three skills are the gateway to practical NumPy.
Position → Range → Condition. First point to one value, then a region, then let a rule choose the region for you.
response = np.array([18, 42, 27, 61, 35, 75])
print(response[2])
print(response[1:4])
print(response[response > 40]) # only slower casesWatch out for views: many basic slices refer to the same underlying data. Editing a slice can therefore affect the original array. Use .copy() when you need an independent copy.
Vectorization means expressing a calculation as an operation on an entire array. NumPy executes many of these operations in optimized compiled code.
Instead of “for each number, do X,” think “apply X to the whole array.” That mental shift is one of the most important NumPy skills.
minutes = np.array([30, 45, 52, 28])
target = 40
variance = minutes - target
status = np.where(minutes <= target, 'On target', 'Late')
print(variance)
print(status)Vectorization shines when the same numerical rule applies across many values. A loop is still valid when every item requires complex branching, external I/O, or object-level logic.
Broadcasting lets NumPy combine arrays with compatible shapes without manually copying the smaller array. It is powerful because it expresses rules by dimension.
A scalar can stretch across every cell. A 1-D row can stretch across rows of a 2-D matrix. Compatibility is checked from the trailing dimensions.
sales = np.array([[100, 120, 90],
[80, 110, 105]])
factors = np.array([1.00, 1.05, 0.98])
adjusted = sales * factors
print(adjusted)When two shapes will not broadcast, reshape deliberately. For example, a per-row factor may need factors[:, None] so NumPy sees a column instead of a flat 1-D vector.
Aggregations compress many values into summaries. The key is understanding axis: which dimension disappears after the calculation?
For a 2-D table: axis=0 aggregates down the rows and returns one result per column; axis=1 aggregates across columns and returns one result per row.
daily = np.array([[22, 30, 35],
[18, 26, 40],
[25, 28, 32]])
print(daily.mean())
print(daily.mean(axis=0)) # per column
print(daily.mean(axis=1)) # per rowIf an axis result surprises you, compare the input shape and output shape. The reduced axis is the one removed unless keepdims=True.
Random numbers support simulation, sampling, testing and synthetic data. Modern NumPy examples use a Generator created with np.random.default_rng().
Random does not have to mean unrepeatable. A fixed seed creates a reproducible sequence — critical for debugging and teaching.
rng = np.random.default_rng(42)
simulated_minutes = rng.normal(loc=35, scale=8, size=1000)
print(simulated_minutes.mean())
print(np.percentile(simulated_minutes, [50, 90, 95]))A seed gives repeatability, not “better randomness.” In analysis, record the seed when reproducibility matters.
The real skill is not memorizing functions. It is recognizing when a problem can be translated into arrays, shapes, masks, vectorized rules, broadcasting and aggregation.
Business question → Numeric array → Clean shape → Vectorized rule → Aggregate → Interpret → Hand off to Pandas / visualization / ML when labels and richer context are needed.
# Mini case: service response minutes by team and day
response = np.array([[32, 35, 40, 28],
[45, 50, 39, 42],
[25, 27, 31, 29]])
target = np.array([35, 40, 30])[:, None]
team_avg = response.mean(axis=1)
within_target = team_avg <= target.ravel()
print(team_avg)
print(within_target)Market skill: explain the transformation, not just the syntax. Employers value analysts who can connect array operations to measurable decisions, quality checks and downstream models.
Open each item only after answering it in your own words.
Shape describes dimensions; size is the total number of elements.
Because basic slicing commonly creates a view that shares the same underlying data.
It lets compatible shapes participate in the same operation without manually duplicating the smaller array.
Write down the input shape, identify the dimension being reduced, and compare the output shape.
A fixed seed makes the pseudo-random results reproducible for teaching, testing and debugging.
| Need | Python List | NumPy | Pandas |
|---|---|---|---|
| Simple general-purpose container | ✅ | ○ | ○ |
| Dense numerical computation | ○ | ✅ | ✅/○ |
| Broadcasting / matrix-style math | ○ | ✅ | ✅/○ |
| Named columns, joins, dates, mixed types | ○ | ○ | ✅ |
| Foundation for scientific / ML stack | ○ | ✅ | ✅ |
The technical concepts in this training follow the NumPy 2.x documentation.
“I can convert numerical business rules into NumPy arrays, validate shapes and dtypes, filter with Boolean masks, vectorize calculations, use broadcasting deliberately, summarize by axis, and create reproducible random simulations.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.