dtype('int64')
dtype('float64')
dtype('bool')
dtype('<U3')
Lecture 08
Every NumPy array stores elements of a single type, its dtype. As with R’s atomic vectors the type is inferred at creation and mixed values are coerced, but NumPy has many more types and each has an explicit size - int8 to int64, uint8 to uint64, float16 to float64, bool, complex128, etc.
The dtype is set when the array is created and does not change. Where an R vector is promoted to fit what is assigned into it, a NumPy array converts the value to its own dtype.
Python has no built-in tabular data structure. pandas (2008) is an established choice - it grew out of NumPy and borrows heavily from R.
Three classes do the heavy lifting:
Series - a 1d typed array with an index (a column)
Index - the labels attached to rows and columns
DataFrame - a collection of Series sharing an index (a table)
A Series is a 1d array with a dtype together with an index of labels - the closest thing in R is a named vector. The difference is that a Series always has an index (integer positions by default) and the index participates in lookup, alignment, and joins rather than being decoration/annotation.
pandas dtypes come from NumPy (int64, float64, bool, datetime64) plus extension dtypes (str, category, Int64, …).
For mixtures such as numbers and strings, pandas falls back to object rather than coercing everything as R does. An object Series still supports elementwise operations, but storing arbitrary Python objects usually costs more memory and computation.
A pandas index can hold integers, strings, etc. A scalar s[key] looks up a label, even when the key is an integer. Use .iloc for positions and .loc for labels.
.loc label slices include both endpoints; .iloc position slices exclude the stop. Integer slices in s[...] are positional, unlike scalar integer lookups.
Arithmetic between Series aligns on index rather than position. The result spans the union of both indexes with NaN for labels missing on one side. Assigning a Series to a column in a DataFrame aligns the same way.
A DataFrame is a dict-like collection of Series sharing a row index, (keys are the column indexes).
species island bill_length_mm ... body_mass_g sex year
0 Adelie Torgersen 39.1 ... 3750.0 male 2007
1 Adelie Torgersen 39.5 ... 3800.0 female 2007
.. ... ... ... ... ... ... ...
342 Chinstrap Dream 50.8 ... 4100.0 male 2009
343 Chinstrap Dream 50.2 ... 3775.0 female 2009
[344 rows x 8 columns]
R has typed NA values each atomic vector type. NumPy has no universal missing-value sentinel across dtypes: floating point arrays can use NaN and datetimes can use NaT, but integers and booleans have no equivalent. With pandas’ default interface, missing integers become float64, while the boolean example below becomes object holding None.
As we’ve seen previously, NaN is not equal to anything, including itself. So == cannot find missing values - isna() (and notna()) recognize NaN, None, NaT, and pd.NA alike. dropna() and fillna() can be used to remove or replace.
Reductions skip missing values by default (skipna=True), the opposite of R’s na.rm = FALSE default..
The extension dtypes Int64, Float64, boolean, and string use pd.NA to represent missing values without changing dtype. Only the numeric names here are capitalized. These types are opt in, via dtype=, dtype_backend= when reading, or convert_dtypes(). Arithmetic and comparisons use three valued logic - pd.NA == 1 is <NA> rather than False.
[]df[...] is overloaded on the type of its argument - a string picks a column, a list of strings a set of columns, and a slice or boolean Series picks rows. There is no two argument df[i, j] form.
.loc and .ilocTwo dimensional subsetting like R’s df[i, j] goes through .loc (labels) and .iloc (positions), each taking [rows, cols]. Label slices include the stop; position slices exclude it. Both accept boolean arrays, but an indexed boolean Series belongs with .loc, not .iloc.
Use & for elementwise AND, | for OR, and ~ for NOT. Parenthesize comparisons; Python’s and / or do not combine Series. .isin() tests membership in a collection.
Selecting a single row returns a Series, and a Series has one dtype - so a heterogeneous row is upcast to object. R keeps a single row as a data frame for exactly this reason.
DataFrames are mutable - df["z"] = ... changes the object in place, and any other name bound to the same object sees the change (reference semantics). This is unlike R’s copy-on-modify, where the same code would leave d untouched. copy() makes an independent data frame, and methods such as assign() return a new one.
Use a single .loc[rows, column] = value assignment on the frame you want to change. Chained assignment such as df["x"][mask] = value cannot update the original frame under copy-on-write.
dplyr’s filter(df, x > 1) works because R captures the unevaluated expression. Python evaluates arguments eagerly and a bare x is a NameError, so pandas offers several workarounds, none of which are quite as convenient.
Explicit - reuse the object,
Callable - a function / lambda
There is no pipe operator, but since nearly every method returns a new DataFrame the equivalent is method chaining. Callables and pd.col() are what let a step refer to columns of the intermediate result. Subsetting with [] can be a step in the chain like any other.
Split-apply-combine works the same way - groupby() returns a DataFrameGroupBy which is then used by subsequent aggregation calls. By default the group keys become the index of the result (as_index=False keeps them as columns, like .by). Named aggregations use (column, function) pairs.
transform() returns a result the same length as its input, broadcasting each group’s value back over that group’s rows - the equivalent of mutate() with .by. Since the result shares the frame’s index it can be used directly in arithmetic or assignment.
Group by two keys and the result has a hierarchical row MultiIndex; aggregate a column two ways with the dictionary/list syntax and the columns get one too. The tidyverse usually keeps group keys and summaries as ordinary columns.
Using pandas and data/flights.parquet, answer the following. Questions 1 and 2 are the core; 3 and 4 are extensions if time allows.
How many flights to LAX did each legacy carrier (AA, UA, DL, US) have in May from JFK, and what was their mean actual air time (air_time)? Report only carriers with matching records.
Which plane (tailnum) has the most flight records from each New York airport? Exclude missing tail numbers and return all ties.
Which five calendar dates (year, month, day) had the lowest mean departure delay? Skip missing delays and break ties by earliest date.
Which flight has the largest arrival delay as a percentage of its actual air time (100 * arr_delay / air_time)? Exclude missing values and nonpositive air times, and return all ties.
| Concept | R / tibble | pandas |
|---|---|---|
| column | atomic vector | Series (values + index) |
| row labels | row.names, rarely used |
Index, central |
| alignment | by position, with recycling | by index label |
| mixed types | coerced to a common type | object dtype |
| missing values | NA in any type |
NaN / None / NaT / pd.NA by dtype |
| column references | data masking | strings, callables, pd.col() |
| mutation | copy-on-modify | in place, copy-on-write for subsets |
| grouped result | keys as columns | keys as (Multi)Index |
| dplyr | pandas |
|---|---|
filter(x > 1) |
query("x > 1"), [pd.col("x") > 1] |
select(x, y) |
[["x", "y"]], filter(regex=) |
mutate(z = x * 2) |
assign(z = pd.col("x") * 2) |
arrange(desc(x)) |
sort_values("x", ascending=False) |
summarize(m = mean(x)) |
agg(m = ("x", "mean")) |
group_by(g) |
groupby("g", as_index=False) |
mutate(m = mean(x), .by = g) |
groupby("g")["x"].transform("mean") |
across(where(is.numeric), f) |
select_dtypes("number").apply(f) |
rename(new = old) |
rename(columns={"old": "new"}) |
A pandas Series is a typed array plus an index, and the index is what makes pandas different from R - lookup is by label, arithmetic and assignment align on labels, and group keys end up in the index. .loc and .iloc exist because [] cannot tell labels from positions.
pandas dtypes are messier than R’s - mixed data becomes object, and missing values are NaN, None, NaT, or pd.NA depending on the dtype, with NaN silently turning integer columns into floats. Use isna(), never ==, and remember that aggregations skip missing values by default.
Python has no data masking, so pandas refers to columns with strings, callables, or pd.col() and chains methods instead of piping. DataFrames are mutable, so aliases matter.
Sta 523 - Fall 2026