Data Structures
in R & Python

Lecture 06

Dr. Colin Rundel

Python classes

Basic syntax

A class statement bundles attributes (data) and methods (functions) into a new type. Methods receive the instance as their first argument, conventionally named self.

class rect:
  """An object representing a rectangle"""
  p1 = (0, 0)
  p2 = (1, 2)

  def area(self):
    return abs((self.p1[0] - self.p2[0]) *
               (self.p1[1] - self.p2[1]))

  def set_p1(self, p1):
    self.p1 = p1
x = rect()
x.area()
2
x.p1
(0, 0)
x.set_p1((1, 1))
x.area()
0
x.p2 = (3, 3)
x.area()
4

__init__

Class construction, i.e. calling the class rect(), creates a new instance of the class its __init__() method.

class rect:
  """An object representing a rectangle"""

  def __init__(self, p1=(0, 0), p2=(1, 1)):
    self.p1 = p1
    self.p2 = p2

  def area(self):
    return abs((self.p1[0] - self.p2[0]) *
               (self.p1[1] - self.p2[1]))
rect().area()
1
rect((0, 0), (3, 3)).area()
9
z = rect(p1=(-1, -1))
z.p1
(-1, -1)
z.p2
(1, 1)

Class vs instance attributes

Attributes assigned in the class body are class attributes shared by every instance (useful for constants), while attributes assigned via self are instance attributes belonging to a specific object.

class die:
  sides = 6

  def __init__(self, value=1):
    self.value = value
d = die(3)
d.value
3
d.sides
6
die.sides
6
die.value
AttributeError: type object 'die' has no attribute 'value'
vars(d)
{'value': 3}
[m for m in dir(d) if not m.startswith("__")]
['sides', 'value']
die.sides = 3
d.sides
3

Mutate or return?

Python methods like sort() and append() modify the object in place and return None, while functions like sorted() return a new object and leave the original alone. R functions essentially always do the latter, so results must be assigned back to be kept.

x = [3, 1, 2]
print( x.sort() )
None
x
[1, 2, 3]
x = [3, 1, 2]
sorted(x)
[1, 2, 3]
x
[3, 1, 2]
x = c(3, 1, 2)
sort(x)
[1] 1 2 3
x
[1] 3 1 2
x = sort(x)
x
[1] 1 2 3

Method chaining

Returning self (or a new object) from a method allows calls to be chained, with each method operating on the result of the previous one. This plays the same role in Python that the pipe does in R.

def set_p1(self, p1):
  self.p1 = p1
  return self

rect.set_p1 = set_p1
( rect()
  .set_p1((-1, -1))
  .area()
)
4
" Hello World ".strip().lower().split()
['hello', 'world']

__str__ and __repr__

Every object has a default string representation, which can be overridden these dunder methods:

rect()
<__main__.rect object at 0x114ceb100>
print(rect())
<__main__.rect object at 0x114d2a8d0>
  • __repr__() - returns an unambiguous representation (ideally valid Python).

  • __str__() - used by print() and str(), falls back to __repr__().

def rect_repr(self):
  return f"rect({self.p1}, {self.p2})"

rect.__repr__ = rect_repr
rect()
rect((0, 0), (1, 1))
print(rect())
rect((0, 0), (1, 1))
[rect(), rect((1, 1), (2, 2))]
[rect((0, 0), (1, 1)), rect((1, 1), (2, 2))]

Other special methods

Operators and built-in functions also dispatch to special methods - this is the mechanism behind the operator overloading we’ve seen previously.

def rect_eq(self, other):
  return (self.p1 == other.p1 and 
          self.p2 == other.p2)

rect.__eq__ = rect_eq

def rect_contains(self, pt):
  return all(
    min(a, b) <= x <= max(a, b)
    for a, b, x in zip(self.p1, self.p2, pt)
  )

rect.__contains__ = rect_contains
rect()
rect((0, 0), (1, 1))
rect() == rect()
True
rect() == rect((1, 1), (2, 2))
False
(0.5, 0.5) in rect()
True
(2, 2) in rect()
False

Inheritance

A class can inherit from another class(es), gaining all of its attributes and methods. The subclass can add new methods and override existing ones - super() gives access to the parent’s version.

class square(rect):
  def __init__(self, p1=(0, 0), l=1):
    if not isinstance(l, (int, float)):
      raise TypeError("l must be a number")
    self.l = l
    super().__init__(p1, (p1[0] + l, p1[1] + l))

  def __repr__(self):
    return f"square({self.p1}, {self.l})"
s = square((1, 1), 2)
s
square((1, 1), 2)
s.p2
(3, 3)
s.area()
4
(2, 2) in s
True
square(l="a")
TypeError: l must be a number

isinstance() vs type()

isinstance() respects inheritance while type() reports the exact class - the former is almost always what you want when validating inputs, but be aware of what type inherits what.

isinstance(s, square)
True
isinstance(s, rect)
True
type(s) is rect
False
isinstance(3, (int, float))
True
isinstance("3", (int, float))
False
isinstance(True, int)
True
type(True) is int
False
bool.__mro__
(<class 'bool'>, <class 'int'>, <class 'object'>)

Tuples

Tuples

Python tuples are heterogeneous, ordered, immutable containers - lists whose contents cannot be changed. R has no tuple type, a list (or an atomic vector) is used instead.

(1, 2, 3)
(1, 2, 3)
(1, True, "abc")
(1, True, 'abc')
type( (1) )
<class 'int'>
type( (1,) )
<class 'tuple'>
x = (1, 2, 3)
x[2] = 5
TypeError: 'tuple' object does not support item assignment
x = (1, [2, 3])
x[1].append(4)
x
(1, [2, 3, 4])

Unpacking

Assigning a sequence to several names at once unpacks it. A name prefixed with * collects any remaining values into a list.

def minmax(x):
    return min(x), max(x)
minmax([3, 1, 2])
(1, 3)
lo, hi = minmax([3, 1, 2])
print(lo, hi)
1 3
x, y = 1, 2
x, y = y, x
x, y
(2, 1)
first, *rest = range(6)
first
0
rest
[1, 2, 3, 4, 5]
pairs = [("a", 1), ("b", 2)]
for i, (k, v) in enumerate(pairs):
    print(i, k, v)
0 a 1
1 b 2

Dictionaries

Dictionaries

Python dicts are mutable containers of key: value pairs, designed for the efficient lookup of a value by its key. The closest (common) R analog is a named list (or a named atomic vector).

{"abc": 123, "def": 456}
{'abc': 123, 'def': 456}
dict([("abc", 123), ("def", 456)])
{'abc': 123, 'def': 456}
dict(abc=123, hello=456)
{'abc': 123, 'hello': 456}
dict(zip(["abc", "def"], [1, 2]))
{'abc': 1, 'def': 2}
list(abc = 123, def = 456)
$abc
[1] 123

$def
[1] 456
c(abc = 123, def = 456)
abc def 
123 456 

Keys must be hashable

Dictionary keys and set elements must be hashable and values can be anything.

Hashable Unhashable
int, float, bool, complex list
str, bytes, range dict
None, frozenset set
tuple if all its elements are hashable tuple containing an unhashable element

Lists are unhashable, so neither a list nor a tuple containing a list can be a key:

{[1]: "bad"}
TypeError: cannot use 'list' as a dict key (unhashable type: 'list')
{(1, [2]): "bad"}
TypeError: cannot use 'tuple' as a dict key (unhashable type: 'list')

You can always coerce a list to a tuple (but this is shallow),

{tuple([1]): "Okay"}
{(1,): 'Okay'}
{(1, tuple([2])): "Okay"}
{(1, (2,)): 'Okay'}

Lookup

Key lookup is via [], if the key is missing a KeyError is raised. Alternatively, use .get() with a default. in checks for the presence of a key.

x = {1: "abc", "y": "hello", (1, 1): 3.14}
x[1]
'abc'
x[(1, 1)]
3.14
x["def"]
KeyError: 'def'
"y" in x
True
"hello" in x
False
print( x.get("def") )
None
x.get("def", 0)
0

Insert, replace, remove

Assigning to d[key] inserts or replaces a value for that key. del and .pop() remove key & value.

x = {1: "abc", "y": "hello"}
x["def"] = -1
x["y"] = "goodbye"
x
{1: 'abc', 'y': 'goodbye', 'def': -1}
del x[1]
x.pop("def")
-1
x["y"] = None
x
{'y': None}
y = list(abc = 123, def = 456)
y$ghi = -1
y$def = "goodbye"
str(y)
List of 3
 $ abc: num 123
 $ def: chr "goodbye"
 $ ghi: num -1
y$abc = NULL
str(y)
List of 2
 $ def: chr "goodbye"
 $ ghi: num -1

Methods & iteration

Iterating over a dict yields its keys - use .items() to get keys and values together.

x = {1: "abc", "y": "hello"}
x.keys()
dict_keys([1, 'y'])
x.values()
dict_values(['abc', 'hello'])
for k, v in x.items():
    print(k, "->", v)
1 -> abc
y -> hello
x.items()
dict_items([(1, 'abc'), ('y', 'hello')])
x | {"y": "goodbye", "w": 0}
{1: 'abc', 'y': 'goodbye', 'w': 0}
x.update({"y": "goodbye", "w": 0})
x
{1: 'abc', 'y': 'goodbye', 'w': 0}

Exercise 1

Using the record from last lecture, complete these tasks in Python:

person = {
  "firstName": "John",
  "lastName": "Smith",
  "age": 25,
  "address": {
    "streetAddress": "21 2nd Street",
    "city": "New York",
    "state": "NY",
    "postalCode": 10021
  },
  "phoneNumber": [
    { "type": "home",
      "number": "212 555-1239" },
    { "type": "fax",
      "number": "646 555-4567" }
  ]
}
  • Extract the fax number.

  • Add a "mobile" phone number and remove age.

  • Write phone(person, phone_type) to return the requested number, or None if absent.

Sets

Sets

A set is a mutable, unordered collection of unique hashable elements - there are no positions, so the primary operation is membership testing with in.

x = {1, 2, 3, 4, 1, 2}
x
{1, 2, 3, 4}
set("mississippi")
{'p', 's', 'm', 'i'}
3 in x
True
x[0]
TypeError: 'set' object is not subscriptable
x.add(9)
x.discard(8)
x.update([7, 8])
x
{1, 2, 3, 4, 7, 8, 9}
x.remove(6)
KeyError: 6
{1, 2, [1, 2]}
TypeError: cannot use 'list' as a set element (unhashable type: 'list')

Set operations

Sets support the usual mathematical operations. R has no set type - any vector can be treated as one via unique() and the set functions.

x = {1, 2, 3}
y = {2, 3, 4}
x | y
{1, 2, 3, 4}
x & y
{2, 3}
x - y
{1}
1 in x
True
x = c(1, 2, 3, 1)
y = c(2, 3, 4)
union(x, y)
[1] 1 2 3 4
intersect(x, y)
[1] 2 3
setdiff(x, y)
[1] 1
1 %in% x
[1] TRUE

Algorithms & data structures

Big-O notation

Big-O notation describes the complexity of an algorithm - how the time (or memory) required grows with the size of the input \(n\). Only the fastest growing term matters (constant factors and lower order terms are dropped), so two algorithms with the same Big-O can still differ greatly in practice.

Since performance depends on the data we usually quote average or worst case complexity, and amortized complexity when an occasional expensive step is paid for by many cheap ones (e.g. growing a list).

Complexity Big-O
Constant O(\(1\))
Logarithmic O(\(\log n\))
Linear O(\(n\))
Quasilinear O(\(n \log n\))
Quadratic O(\(n^2\))
Exponential O(\(C^n\))

Under the hood

The containers in both languages are built from a handful of basic structures - knowing which is which explains what each container is good, and bad, at.

Structure Layout R Python
Array contiguous block of same-sized values materialized numeric vector contiguous numeric ndarray
Array of pointers contiguous block of references to objects list list, tuple
Linked list nodes that each point to their neighbors pairlist deque (linked blocks)
Hash table array indexed by hash(key) environment dict, set

Vectors

Materialized numeric vectors in R and contiguous numeric NumPy arrays store fixed-size values together, giving O(1) element access. NumPy views can have gaps or reversed strides.

object.size(numeric(1e6))
8000048 bytes
object.size(integer(1e6))
4000048 bytes
np.zeros(1_000_000).nbytes
8000000
np.zeros(1_000_000, dtype="int64").nbytes
8000000
np.zeros(1_000_000, dtype="int32").nbytes
4000000

Growing a vector

Repeatedly appending with c(x, i) copies the existing elements each time, giving O(\(n^2\)) work overall.

n = 5e4
x = c()
system.time(
  for (i in seq_len(n)) x = c(x, i)
)
   user  system elapsed 
  1.028   0.016   1.044 

Indexed growth can reuse spare capacity, but still needs occasional copying. Preallocation avoids repeated growth.

x = numeric()
system.time(
  for (i in seq_len(n)) x[i] = i
)
   user  system elapsed 
  0.004   0.000   0.004 
x = numeric(n)
system.time(
  for (i in seq_len(n)) x[i] = i
)
   user  system elapsed 
  0.001   0.000   0.001 

Array (Vector)

Generic vectors

R and Python lists store references to objects. Elements can have different types and sizes, while indexing remains O(1). References and object headers add memory overhead.

  • Numeric vector - values stored contiguously:

    [ 1.0 | 2.0 | 3.0 ]
  • List - references stored contiguously

    [  P  |  P  |  P  ]
       ↓     ↓     ↓
       1   "abc"  TRUE

deques

A deque (double-ended queue) supports O(1) additions and removals at either end. CPython stores it as a doubly linked chain of fixed-size blocks. Indexing is O(1) near either end and O(n) in the middle.

from collections import deque
x = deque(range(3))
x.appendleft(-1)
x
deque([-1, 0, 1, 2])
x.popleft()
-1
x
deque([0, 1, 2])
x = deque(range(3), maxlen=4)
x.append(10)
x.append(11)
x
deque([1, 2, 10, 11], maxlen=4)
x.appendleft(-1)
x
deque([-1, 1, 2, 10], maxlen=4)

Linked list

Hashing

A hash table is an array of buckets - a key’s hash() picks its bucket, so lookup, insertion and deletion are O(1) on average with no scanning. Two keys landing in the same bucket is a collision, and the table is resized as it fills to keep collisions rare.

hash(1), hash(1.0), hash(True)
(1, 1, 1)
hash(2.5)
1152921504606846978
hash((1, 2))
-3550055125485641917
keys = (1, 9, 17, 2.5)
[hash(k) % 8 for k in keys]
[1, 1, 1, 2]
h = new.env(size = 8L)
for (k in letters[1:5]) h[[k]] = k
str(env.profile(h))
List of 3
 $ size   : int 8
 $ nchains: int 5
 $ counts : int [1:8] 0 1 1 1 1 1 0 0
for (k in letters[6:20]) h[[k]] = k
env.profile(h)$size
[1] 25

Hash table

Operation costs

Operation list deque dict set
Copy .copy(): O(n) .copy(): O(n) .copy(): O(n) .copy(): O(n)
Get / set by index or key s[i]: O(1) s[i]: O(n) d[k]: O(1) —
Append / add .append(v): O(1)* .append(v): O(1) d[k] = v: O(1)* .add(v): O(1)*
Pop from end / remove .pop(): O(1)* .pop(): O(1) .pop(k): O(1) .remove(v): O(1)
Insert at front .insert(0, v): O(n) .appendleft(v): O(1) — —
Remove from front .pop(0): O(n) .popleft(): O(1) — —
Insert / delete in middle O(n) O(n) — —
Membership v in s: O(n) v in s: O(n) k in d: O(1) v in s: O(1)

Exercise 2

For each of the following, suggest a data structure in Python and R. Explain which operations your choice makes efficient, and state any assumptions.

  • A collection of 100 integers used repeatedly in elementwise arithmetic.

  • A queue (first in, first out) of customer records.

  • A stack (first in, last out) of customer records.

  • A count of word occurrences within a document.

  • The heights of the bars in a histogram with even bin widths.

  • Checking whether each of a million words is in a list of 50 stop words.

Measuring vector growth

Reasoning about complexity tells you what to expect, measuring confirms it

grow_concat = function(n) {
  x = c()
  for (i in seq_len(n)) x = c(x, i)
}
grow_index = function(n) {
  x = numeric()
  for (i in seq_len(n)) x[i] = i
}
grow_prealloc = function(n) {
  x = numeric(n)
  for (i in seq_len(n)) x[i] = i
}
res = bench::press(
  n = c(1e3, 2.5e3, 5e3, 1e4, 2.5e4, 5e4),
  bench::mark(
    concat   = grow_concat(n),
    index    = grow_index(n),
    prealloc = grow_prealloc(n),
    check = FALSE
  )
)

Comparing R & Python

Container summary

Python Properties R
list ordered, mutable, heterogeneous list()
tuple ordered, immutable, heterogeneous list() or atomic vector (copy on modify)
dict key lookup, mutable, insertion ordered named list(), new.env() for hashing
set unique, unordered, mutable vector + unique(), union(), %in%, …
deque fast at both ends none (use a list or vector)
range lazy integer sequence, immutable seq_len(), : (often stored compactly via ALTREP)
NumPy ndarray homogeneous, vectorized atomic vector, matrix, array

Takeaways

  • Tuples are immutable sequences—use them for fixed groups of values and unpacking.

  • Dictionaries map hashable keys to values, with O(1) lookup on average. R offers named lists and hash-based environments.

  • Sets provide O(1) membership on average and set algebra. In R, use %in%, union(), intersect(), and setdiff().

  • Shallow copies share nested objects. Use deepcopy() when nested objects need independent mutation.

  • Mutable defaults are shared between calls. Use None and create the object inside the function.

  • Choose containers for their operations: deques for queues, sets for membership, and preallocated R vectors when the size is known. Measure when performance matters.