[1] 1 4 9 16 25
[1] 1.000000 1.414214 1.732051 2.000000 2.236068
[1] 11 12 13 14 15
Lecture 04
Vectorized code applies one operation to many values without writing an explicit loop. R’s arithmetic, comparison, and element-wise logical operators are vectorized.
When vector lengths differ, R repeats the shorter vector to match the longer one.
Zero-length vectors contain no values; NULL also has length zero. For vectorized operations, the presence of a zero-length vector means the result will also have length 0.
Python’s list operators manipulate the container, not the values. Element-wise arithmetic and comparisons instead require explicit iteration: a for loop or, more idiomatically, a comprehension.
A comprehension combines an expression with a for clause to construct a list. A trailing if filters values; a conditional expression transforms every value.
Comprehensions can be nested. Additional for clauses act like nested loops and produce one flat list, while nested comprehensions build a list of lists.
NumPy is a third-party Python package for numerical computing. Its main data structure is the homogeneous, multidimensional ndarray, with operations designed to work efficiently across entire arrays (i.e. vectorization).
NumPy arrays support the same style of vectorized arithmetic, comparison, and logical operations as R’s atomic vectors, with a similar performance advantage over explicit for loops.
np.where() is NumPy’s equivalent to R’s ifelse(); both construct a new vector by choosing between two values element-wise based on a condition.
R uses 1-based indexing; Python uses 0-based indexing when subsetting by non-negative integers.
A slice has the form start:stop:step; start is included and stop is excluded. Any component may be omitted.
R accepts a vector of positions inside [ ]. A base Python list does not: a slice can select a regular range, but arbitrary positions require a comprehension.
Unlike base Python lists, NumPy arrays accept a list or integer array of positions. As with list indexing, negative positions count backward from the end and can be mixed with positive positions.
Out-of-bounds indexing fails in Python, while R returns a typed missing value for a position that does not exist. Python slices clip each endpoint to the nearest valid boundary.
We have already used positive integers to select positions and negative integers to exclude positions. R’s [ operator supports four additional index forms:
logicals - select by TRUE positions
characters - select by names
zero - select nothing
empty index - select everything
Logical subsetting keeps values corresponding to TRUE.
Most logical indexes are created by vectorized logical expressions, which return one logical value for each element. Such a vector is often called a mask.
NumPy also supports using Boolean arrays for subsetting - positions where the index is True are kept and those with False are discarded. Boolean ndarrays, lists of Booleans, and the masks produced by vectorized comparisons all work.
However, while R recycles logical vectors, NumPy requires a Boolean index to match the dimension it selects.
NumPy overloads &, |, and ~ as element-wise logical operators.
Parenthesize every comparison because these operators have higher precedence than comparison operators.
Names allow values to be selected independently of their positions.
In R, an empty index selects everything and a zero index selects nothing. This differs from Python, where 0 is the first position and : is used to select everything.
Subset syntax can appear on the left side of an assignment in R, base Python, and NumPy. A base Python list supports assignment to a single position or a slice, while R and NumPy also allow assignment via logical / Boolean and integer indexes.
Basic slicing always returns a view that shares data with the original array. Advanced indexing returns a copy.
Without running the code, determine the result (or error) of each expression.
Both are homogeneous, multidimensional containers designed for vectorized computation.
R matrices use column-major ordering - a sequence of values fills the matrix down each column. NumPy arrays default to row-major ordering - values fill across each row.
Dimensions are separated by commas in both languages; ranges select rectangular regions.
Selecting a single row or column with a scalar index removes that dimension. Use drop = FALSE in R or a length-one slice in NumPy to preserve a two-dimensional result.
An integer list can select along either dimension. When both dimensions receive integer lists, NumPy pairs them element by element instead of crossing them.
The paired expression selects (row 0, column 1) and (row 2, column 3), not a \(2 \times 2\) rectangle; np.ix_() produces the rectangle.
A reduction combines many values into fewer values. NumPy’s axis argument specifies which dimension is collapsed.
R recycles by comparing lengths; NumPy broadcasts by comparing shapes.
Shapes are compared from the rightmost dimension to the left.
Each pair of dimensions is compatible when:
they are equal, or
one of them is 1.
An array with fewer dimensions is treated as if it had leading dimensions of size 1.
For one-dimensional arrays, sizes must match or one of the arrays must have size 1.
A one-dimensional array aligns with the trailing (column) dimension, so it is applied to every row.
To align with the row dimension instead, give the vector an explicit singleton column dimension.
Broadcasting two singleton dimensions produces every pairwise combination.
Broadcasting makes it possible to transform every column using its own mean and standard deviation.
Valid shapes do not guarantee the intended calculation: both expressions below run, but only the second subtracts each row’s own mean.
For each pair of NumPy shapes, determine whether broadcasting succeeds and, if so, the result shape.
(128, 128, 3) and (3,)
(8, 1, 6, 1) and (7, 1, 5)
(2, 1) and (8, 4, 3)
(3, 1) and (15, 3, 5)
(3,) and (4,)
Then write a vectorized NumPy expression that subtracts the mean of each row from a matrix x with shape (100, 5).
| Concept | R | Python / NumPy |
|---|---|---|
| element-wise arithmetic | built into atomic vectors | NumPy arrays |
| element-wise comparisons | built into atomic vectors | NumPy arrays |
| logical operators | &, |, ! |
&, |, ~ |
| element-wise choice | ifelse() |
np.where() |
| scalar repetition | recycling | broadcasting |
| unequal sizes | repeat by total length | compare shapes from the right |
| explicit iteration | for, later lapply() / purrr |
comprehension, for |
| reduction | sum(), mean(), rowMeans() |
.sum(), .mean(axis=...) |
| Goal | R | Python / NumPy |
|---|---|---|
| first element | x[1] |
x[0] |
| last element | x[length(x)] |
x[-1] |
| exclude first | x[-1] |
x[1:] |
| regular range | x[2:5] |
x[1:5] |
| arbitrary positions | x[c(1, 3)] |
x[[0, 2]] (NumPy) |
| Boolean filter | x[x > 0] |
x[x > 0] (NumPy) |
| all rows, second column | x[, 2] |
x[:, 1] |
| independent copy | automatic (copy-on-modify) | x.copy() (NumPy) |
R atomic vectors and NumPy arrays support element-wise operations; Python lists require explicit iteration or comprehensions.
R indexes from 1 and uses negative indexes for exclusion; Python indexes from 0, counts backward with negative indexes.
R’s [ supports integer, logical, and name-based selection. NumPy adds integer-array and Boolean indexing to Python’s usual indexing syntax.
Basic NumPy slices are always views; advanced indexing is a copy. Make copying intentional when the result will be modified.
R recycling is based on vector length; NumPy broadcasting is based on compatible shapes.
Sta 523 - Fall 2026