Lecture 06
A class statement bundles attributes (data) and methods (functions) into a new type. Methods receive the instance as their first argument, conventionally named self.
__init__Class construction, i.e. calling the class rect(), creates a new instance of the class its __init__() method.
Attributes assigned in the class body are class attributes shared by every instance (useful for constants), while attributes assigned via self are instance attributes belonging to a specific object.
Python methods like sort() and append() modify the object in place and return None, while functions like sorted() return a new object and leave the original alone. R functions essentially always do the latter, so results must be assigned back to be kept.
Returning self (or a new object) from a method allows calls to be chained, with each method operating on the result of the previous one. This plays the same role in Python that the pipe does in R.
__str__ and __repr__Every object has a default string representation, which can be overridden these dunder methods:
__repr__() - returns an unambiguous representation (ideally valid Python).
__str__() - used by print() and str(), falls back to __repr__().
Operators and built-in functions also dispatch to special methods - this is the mechanism behind the operator overloading we’ve seen previously.
A class can inherit from another class(es), gaining all of its attributes and methods. The subclass can add new methods and override existing ones - super() gives access to the parent’s version.
isinstance() vs type()isinstance() respects inheritance while type() reports the exact class - the former is almost always what you want when validating inputs, but be aware of what type inherits what.
Python tuples are heterogeneous, ordered, immutable containers - lists whose contents cannot be changed. R has no tuple type, a list (or an atomic vector) is used instead.
Assigning a sequence to several names at once unpacks it. A name prefixed with * collects any remaining values into a list.
Python dicts are mutable containers of key: value pairs, designed for the efficient lookup of a value by its key. The closest (common) R analog is a named list (or a named atomic vector).
Dictionary keys and set elements must be hashable and values can be anything.
| Hashable | Unhashable |
|---|---|
int, float, bool, complex |
list |
str, bytes, range |
dict |
None, frozenset |
set |
tuple if all its elements are hashable |
tuple containing an unhashable element |
Lists are unhashable, so neither a list nor a tuple containing a list can be a key:
Key lookup is via [], if the key is missing a KeyError is raised. Alternatively, use .get() with a default. in checks for the presence of a key.
Assigning to d[key] inserts or replaces a value for that key. del and .pop() remove key & value.
{1: 'abc', 'y': 'goodbye', 'def': -1}
Iterating over a dict yields its keys - use .items() to get keys and values together.
Using the record from last lecture, complete these tasks in Python:
Extract the fax number.
Add a "mobile" phone number and remove age.
Write phone(person, phone_type) to return the requested number, or None if absent.
A set is a mutable, unordered collection of unique hashable elements - there are no positions, so the primary operation is membership testing with in.
Sets support the usual mathematical operations. R has no set type - any vector can be treated as one via unique() and the set functions.
Big-O notation describes the complexity of an algorithm - how the time (or memory) required grows with the size of the input \(n\). Only the fastest growing term matters (constant factors and lower order terms are dropped), so two algorithms with the same Big-O can still differ greatly in practice.
Since performance depends on the data we usually quote average or worst case complexity, and amortized complexity when an occasional expensive step is paid for by many cheap ones (e.g. growing a list).
| Complexity | Big-O |
|---|---|
| Constant | O(\(1\)) |
| Logarithmic | O(\(\log n\)) |
| Linear | O(\(n\)) |
| Quasilinear | O(\(n \log n\)) |
| Quadratic | O(\(n^2\)) |
| Exponential | O(\(C^n\)) |
The containers in both languages are built from a handful of basic structures - knowing which is which explains what each container is good, and bad, at.
| Structure | Layout | R | Python |
|---|---|---|---|
| Array | contiguous block of same-sized values | materialized numeric vector | contiguous numeric ndarray |
| Array of pointers | contiguous block of references to objects | list | list, tuple |
| Linked list | nodes that each point to their neighbors | pairlist | deque (linked blocks) |
| Hash table | array indexed by hash(key) |
environment | dict, set |
Materialized numeric vectors in R and contiguous numeric NumPy arrays store fixed-size values together, giving O(1) element access. NumPy views can have gaps or reversed strides.
Repeatedly appending with c(x, i) copies the existing elements each time, giving O(\(n^2\)) work overall.
Indexed growth can reuse spare capacity, but still needs occasional copying. Preallocation avoids repeated growth.
R and Python lists store references to objects. Elements can have different types and sizes, while indexing remains O(1). References and object headers add memory overhead.
Numeric vector - values stored contiguously:
[ 1.0 | 2.0 | 3.0 ]
List - references stored contiguously
[ P | P | P ]
↓ ↓ ↓
1 "abc" TRUE
A deque (double-ended queue) supports O(1) additions and removals at either end. CPython stores it as a doubly linked chain of fixed-size blocks. Indexing is O(1) near either end and O(n) in the middle.
A hash table is an array of buckets - a key’s hash() picks its bucket, so lookup, insertion and deletion are O(1) on average with no scanning. Two keys landing in the same bucket is a collision, and the table is resized as it fills to keep collisions rare.
(1, 1, 1)
1152921504606846978
-3550055125485641917
| Operation | list |
deque |
dict |
set |
|---|---|---|---|---|
| Copy | .copy(): O(n) |
.copy(): O(n) |
.copy(): O(n) |
.copy(): O(n) |
| Get / set by index or key | s[i]: O(1) |
s[i]: O(n) |
d[k]: O(1) |
— |
| Append / add | .append(v): O(1)* |
.append(v): O(1) |
d[k] = v: O(1)* |
.add(v): O(1)* |
| Pop from end / remove | .pop(): O(1)* |
.pop(): O(1) |
.pop(k): O(1) |
.remove(v): O(1) |
| Insert at front | .insert(0, v): O(n) |
.appendleft(v): O(1) |
— | — |
| Remove from front | .pop(0): O(n) |
.popleft(): O(1) |
— | — |
| Insert / delete in middle | O(n) | O(n) | — | — |
| Membership | v in s: O(n) |
v in s: O(n) |
k in d: O(1) |
v in s: O(1) |
For each of the following, suggest a data structure in Python and R. Explain which operations your choice makes efficient, and state any assumptions.
A collection of 100 integers used repeatedly in elementwise arithmetic.
A queue (first in, first out) of customer records.
A stack (first in, last out) of customer records.
A count of word occurrences within a document.
The heights of the bars in a histogram with even bin widths.
Checking whether each of a million words is in a list of 50 stop words.
Reasoning about complexity tells you what to expect, measuring confirms it
| Python | Properties | R |
|---|---|---|
list |
ordered, mutable, heterogeneous | list() |
tuple |
ordered, immutable, heterogeneous | list() or atomic vector (copy on modify) |
dict |
key lookup, mutable, insertion ordered | named list(), new.env() for hashing |
set |
unique, unordered, mutable | vector + unique(), union(), %in%, … |
deque |
fast at both ends | none (use a list or vector) |
range |
lazy integer sequence, immutable | seq_len(), : (often stored compactly via ALTREP) |
NumPy ndarray |
homogeneous, vectorized | atomic vector, matrix, array |
Tuples are immutable sequences—use them for fixed groups of values and unpacking.
Dictionaries map hashable keys to values, with O(1) lookup on average. R offers named lists and hash-based environments.
Sets provide O(1) membership on average and set algebra. In R, use %in%, union(), intersect(), and setdiff().
Shallow copies share nested objects. Use deepcopy() when nested objects need independent mutation.
Mutable defaults are shared between calls. Use None and create the object inside the function.
Choose containers for their operations: deques for queues, sets for membership, and preallocated R vectors when the size is known. Measure when performance matters.
Sta 523 - Fall 2026