---
title: "Data Structures<br/>in R & Python"
subtitle: "Lecture 06"
author: "Dr. Colin Rundel"
footer: "Sta 523 - Fall 2026"
format:
  revealjs:
    theme: slides.scss
    transition: fade
    slide-number: true
    self-contained: true
execute:
  echo: true
  warning: true
engine: knitr
---


```{r setup}
#| message: false
#| warning: false
#| include: false
options(
  width = 80
)
```

```{python py_setup}
#| include: false
import numpy as np
```

# Python classes

## Basic syntax

A `class` statement bundles attributes (data) and methods (functions) into a new type. Methods receive the instance as their first argument, conventionally named `self`.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
class rect:
  """An object representing a rectangle"""
  p1 = (0, 0)
  p2 = (1, 2)

  def area(self):
    return abs((self.p1[0] - self.p2[0]) *
               (self.p1[1] - self.p2[1]))

  def set_p1(self, p1):
    self.p1 = p1
```
:::

::: {.column width='50%' .fragment}
```{python}
x = rect()
x.area()
x.p1
```
```{python}
x.set_p1((1, 1))
x.area()
```
```{python}
x.p2 = (3, 3)
x.area()
```
:::
::::



## `__init__`

Class construction, i.e. calling the class `rect()`, creates a new instance of the class its `__init__()` method.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
class rect:
  """An object representing a rectangle"""

  def __init__(self, p1=(0, 0), p2=(1, 1)):
    self.p1 = p1
    self.p2 = p2

  def area(self):
    return abs((self.p1[0] - self.p2[0]) *
               (self.p1[1] - self.p2[1]))
```
:::

::: {.column width='50%' .fragment}
```{python}
rect().area()
rect((0, 0), (3, 3)).area()
```
```{python}
z = rect(p1=(-1, -1))
z.p1
z.p2
```
:::
::::

::: {.aside}
Methods whose names start and end with double underscores (`__init__`, `__repr__`, ...) are called "dunder" methods - Python calls them implicitly in specific situations.
:::


## Class vs instance attributes

Attributes assigned in the class body are class attributes shared by every instance (useful for constants), while attributes assigned via `self` are instance attributes belonging to a specific object.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
class die:
  sides = 6

  def __init__(self, value=1):
    self.value = value
```
```{python}
d = die(3)
d.value
d.sides
die.sides
```
:::

::: {.column width='50%' .fragment}
```{python}
#| error: true
die.value
```
```{python}
vars(d)
[m for m in dir(d) if not m.startswith("__")]
```
```{python}
die.sides = 3
d.sides
```
:::
::::

::: {.aside}
`dir()` lists all attributes and methods of an object (including inherited dunder methods), while `vars()` returns just the instance attributes as a `dict`.
:::


## Mutate or return?

::: {.medium}
Python methods like `sort()` and `append()` modify the object in place and return `None`, while functions like `sorted()` return a new object and leave the original alone. R functions essentially always do the latter, so results must be assigned back to be kept.
:::

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
x = [3, 1, 2]
print( x.sort() )
x
```

::: {.fragment}
```{python}
x = [3, 1, 2]
sorted(x)
x
```
:::
:::

::: {.column width='50%' .fragment}
```{r}
x = c(3, 1, 2)
sort(x)
x
```
```{r}
x = sort(x)
x
```
:::
::::


## Method chaining

::: {.medium}
Returning `self` (or a new object) from a method allows calls to be chained, with each method operating on the result of the previous one. This plays the same role in Python that the pipe does in R.
:::

::: {.xsmall}
```{python}
def set_p1(self, p1):
  self.p1 = p1
  return self

rect.set_p1 = set_p1
```
:::

. . .

::: {.xsmall}
```{python}
( rect()
  .set_p1((-1, -1))
  .area()
)
```
:::

. . .

::: {.xsmall}
```{python}
" Hello World ".strip().lower().split()
```
:::


::: {.aside}
Chaining in libraries like pandas works the same way, though most of those methods return a new object rather than modifying `self` in place.
:::


## `__str__` and `__repr__`

Every object has a default string representation, which can be overridden these dunder methods:

::: {.xsmall}
```{python}
rect()
print(rect())
```
:::

. . .

::: {.medium}
* `__repr__()` - returns an unambiguous representation (ideally valid Python).

* `__str__()` - used by `print()` and `str()`, falls back to `__repr__()`.
:::

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
def rect_repr(self):
  return f"rect({self.p1}, {self.p2})"

rect.__repr__ = rect_repr
```
```{python}
rect()
```
:::

::: {.column width='50%'}
```{python}
print(rect())
[rect(), rect((1, 1), (2, 2))]
```
:::
::::


## Other special methods

Operators and built-in functions also dispatch to special methods - this is the mechanism behind the operator overloading we've seen previously.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
def rect_eq(self, other):
  return (self.p1 == other.p1 and 
          self.p2 == other.p2)

rect.__eq__ = rect_eq

def rect_contains(self, pt):
  return all(
    min(a, b) <= x <= max(a, b)
    for a, b, x in zip(self.p1, self.p2, pt)
  )

rect.__contains__ = rect_contains
```
:::

::: {.column width='50%' .fragment}
```{python}
rect()
rect() == rect()
rect() == rect((1, 1), (2, 2))
```
```{python}
(0.5, 0.5) in rect()
(2, 2) in rect()
```
:::
::::

::: {.aside}
See the [data model](https://docs.python.org/3/reference/datamodel.html#special-method-names) docs for the full list, including `__eq__`, `__len__`, `__getitem__`, `__add__`, and `__lt__`.
:::


## Inheritance

A class can inherit from another class(es), gaining all of its attributes and methods. The subclass can add new methods and override existing ones - `super()` gives access to the parent's version.

:::: {.columns .xsmall}
::: {.column width='60%'}
```{python}
class square(rect):
  def __init__(self, p1=(0, 0), l=1):
    if not isinstance(l, (int, float)):
      raise TypeError("l must be a number")
    self.l = l
    super().__init__(p1, (p1[0] + l, p1[1] + l))

  def __repr__(self):
    return f"square({self.p1}, {self.l})"
```
:::

::: {.column width='40%' .fragment}
```{python}
s = square((1, 1), 2)
s
s.p2
s.area()
(2, 2) in s
```
```{python}
#| error: true
square(l="a")
```
:::
::::



## `isinstance()` vs `type()`

`isinstance()` respects inheritance while `type()` reports the exact class - the former is almost always what you want when validating inputs, but be aware of what type inherits what.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
isinstance(s, square)
isinstance(s, rect)
type(s) is rect
```
```{python}
isinstance(3, (int, float))
isinstance("3", (int, float))
```
:::

::: {.column width='50%' .fragment}
```{python}
isinstance(True, int)
type(True) is int
bool.__mro__
```
:::
::::


# Tuples

## Tuples

Python tuples are *heterogeneous*, *ordered*, **immutable** containers - lists whose contents cannot be changed. R has no tuple type, a list (or an atomic vector) is used instead.

:::: {.columns .small}
::: {.column width='50%'}
```{python}
(1, 2, 3)
(1, True, "abc")
```
```{python}
type( (1) )
type( (1,) )
```
:::

::: {.column width='50%' .fragment}
```{python}
#| error: true
x = (1, 2, 3)
x[2] = 5
```
```{python}
x = (1, [2, 3])
x[1].append(4)
x
```
:::
::::

::: {.aside}
It is the *comma* that creates a tuple, not the parentheses - `(1)` is just the number 1 and a one element tuple must be written `(1,)`. 

Immutability is *shallow*, a mutable object stored in a tuple can still be modified in place.
:::


## Unpacking

Assigning a sequence to several names at once *unpacks* it. A name prefixed with `*` collects any remaining values into a list.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
def minmax(x):
    return min(x), max(x)
minmax([3, 1, 2])
lo, hi = minmax([3, 1, 2])
print(lo, hi)
```
```{python}
x, y = 1, 2
x, y = y, x
x, y
```
:::

::: {.column width='50%' .fragment}
```{python}
first, *rest = range(6)
first
rest
```
```{python}
pairs = [("a", 1), ("b", 2)]
for i, (k, v) in enumerate(pairs):
    print(i, k, v)
```
:::
::::

::: {.aside}
The number of names must match the length of the sequence (`x, y = [1, 2, 3]` is a `ValueError`) unless one of them is starred.
:::


# Dictionaries

## Dictionaries

Python `dict`s are *mutable* containers of `key: value` pairs, designed for the efficient lookup of a value by its key. The closest (common) R analog is a named list (or a named atomic vector).

:::: {.columns .small}
::: {.column width='50%'}
```{python}
{"abc": 123, "def": 456}
dict([("abc", 123), ("def", 456)])
dict(abc=123, hello=456)
dict(zip(["abc", "def"], [1, 2]))
```
:::

::: {.column width='50%'}
```{r}
list(abc = 123, def = 456)
```
```{r}
c(abc = 123, def = 456)
```
:::
::::

::: {.aside}
Since Python 3.7 dictionaries preserve insertion order. The keyword form of `dict()` only works for keys that are valid Python names (e.g. `def` is a reserved word).
:::


## Keys must be hashable

Dictionary keys and set elements must be hashable and values can be anything.

::: {.small}
| Hashable                                      | Unhashable                                  |
|:----------------------------------------------|:--------------------------------------------|
| `int`, `float`, `bool`, `complex`             | `list`                                      |
| `str`, `bytes`, `range`                       | `dict`                                      |
| `None`, `frozenset`                           | `set`                                       |
| `tuple` if all its elements are hashable      | `tuple` containing an unhashable element     |
:::

. . .

Lists are unhashable, so neither a list nor a tuple containing a list can be a key:

:::: {.columns .small}
::: {.column width='50%'}
```{python}
#| error: true
{[1]: "bad"}
```
:::

::: {.column width='50%'}
```{python}
#| error: true
{(1, [2]): "bad"}
```
:::
::::

. . .

You can always coerce a list to a tuple (but this is shallow),

:::: {.columns .small}
::: {.column width='50%'}
```{python}
#| error: true
{tuple([1]): "Okay"}
```
:::

::: {.column width='50%'}
```{python}
#| error: true
{(1, tuple([2])): "Okay"}
```
:::
::::


## Lookup

Key lookup is via `[]`, if the key is missing a `KeyError` is raised. Alternatively, use `.get()` with a default. `in` checks for the presence of a key.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
#| error: true
x = {1: "abc", "y": "hello", (1, 1): 3.14}
x[1]
x[(1, 1)]
x["def"]
```
:::

::: {.column width='50%' .fragment}
```{python}
"y" in x
"hello" in x
print( x.get("def") )
x.get("def", 0)
```
:::
::::


## Insert, replace, remove

Assigning to `d[key]` inserts or replaces a value for that key. `del` and `.pop()` remove key & value.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
x = {1: "abc", "y": "hello"}
x["def"] = -1
x["y"] = "goodbye"
x
```
```{python}
del x[1]
x.pop("def")
x["y"] = None
x
```
:::

::: {.column width='50%'}
```{r}
y = list(abc = 123, def = 456)
y$ghi = -1
y$def = "goodbye"
str(y)
```
```{r}
y$abc = NULL
str(y)
```
:::
::::

::: {.aside}
Unlike R - assigning `None` does *not* remove a key.
:::

## Methods & iteration

Iterating over a `dict` yields its keys - use `.items()` to get keys and values together.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
x = {1: "abc", "y": "hello"}
```
```{python}
x.keys()
x.values()
```
```{python}
for k, v in x.items():
    print(k, "->", v)
x.items()
```
:::

::: {.column width='50%' .fragment}
```{python}
x | {"y": "goodbye", "w": 0}
```
```{python}
x.update({"y": "goodbye", "w": 0})
x
```
:::
::::

::: {.aside}
`.keys()`, `.values()`, and `.items()` return dynamic views of the dictionary rather than copies - see the [docs](https://docs.python.org/3/library/stdtypes.html#dictionary-view-objects). 

`|` returns a *new* merged dictionary (right side wins), `.update()` merges in place.
:::

## Exercise 1

Using the record from last lecture, complete these tasks in Python:

:::: {.columns}
::: {.column width='55%' .xsmall}

```python
person = {
  "firstName": "John",
  "lastName": "Smith",
  "age": 25,
  "address": {
    "streetAddress": "21 2nd Street",
    "city": "New York",
    "state": "NY",
    "postalCode": 10021
  },
  "phoneNumber": [
    { "type": "home",
      "number": "212 555-1239" },
    { "type": "fax",
      "number": "646 555-4567" }
  ]
}
```
:::

::: {.column width='45%' .medium}
* Extract the fax number.

* Add a `"mobile"` phone number and remove `age`.

* Write `phone(person, phone_type)` to return the requested number, or `None` if absent.
:::
::::


::: {.aside}
JSON uses `true`, `false`, and `null` where Python uses `True`, `False`, and `None`.
:::


# Sets

## Sets

A `set` is a *mutable*, *unordered* collection of **unique** hashable elements - there are no positions, so the primary operation is membership testing with `in`.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
x = {1, 2, 3, 4, 1, 2}
x
set("mississippi")
```
```{python}
#| error: true
3 in x
x[0]
```
:::

::: {.column width='50%' .fragment}
```{python}
x.add(9)
x.discard(8)
x.update([7, 8])
x
```
```{python}
#| error: true
x.remove(6)
```
```{python}
#| error: true
{1, 2, [1, 2]}
```
:::
::::

::: {.aside}
`{}` creates an empty *dictionary*, an empty set must be written `set()`. `remove()` raises a `KeyError` if the element is absent while `discard()` does not, and iteration order is not guaranteed (use `sorted(x)` for a stable order).
:::


## Set operations

Sets support the usual mathematical operations. R has no set type - any vector can be treated as one via `unique()` and the set functions.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
x = {1, 2, 3}
y = {2, 3, 4}
```
```{python}
x | y
x & y
x - y
1 in x
```
:::

::: {.column width='50%'}
```{r}
x = c(1, 2, 3, 1)
y = c(2, 3, 4)
```
```{r}
union(x, y)
intersect(x, y)
setdiff(x, y)
1 %in% x
```
:::
::::

::: {.aside}
Python's set methods (`x.union(y)`, `x.intersection(y)`, ...) accept any iterable, the operators require sets on both sides. R's set functions return unique values in order of first appearance.
:::



# Algorithms & data structures


## Big-O notation

::: {.medium}
Big-O notation describes the *complexity* of an algorithm - how the time (or memory) required grows with the size of the input $n$. Only the fastest growing term matters (constant factors and lower order terms are dropped), so two algorithms with the same Big-O can still differ greatly in practice.

Since performance depends on the data we usually quote *average* or *worst case* complexity, and *amortized* complexity when an occasional expensive step is paid for by many cheap ones (e.g. growing a list).
:::

::: {.small .center}
| Complexity  | Big-O           |
|-------------|-----------------|
| Constant    | O($1$)          |
| Logarithmic | O($\log n$)     |
| Linear      | O($n$)          |
| Quasilinear | O($n \log n$)   |
| Quadratic   | O($n^2$)        |
| Exponential | O($C^n$)        |
:::


## Under the hood

The containers in both languages are built from a handful of basic structures - knowing which is which explains what each container is good, and bad, at.

::: {.small}
| Structure              | Layout                                              | R                                  | Python              |
|:-----------------------|:----------------------------------------------------|:-----------------------------------|:--------------------|
| Array                  | contiguous block of same-sized values               | materialized numeric vector        | contiguous numeric `ndarray` |
| Array of pointers      | contiguous block of *references* to objects         | list                               | `list`, `tuple`     |
| Linked list            | nodes that each point to their neighbors            | pairlist                           | `deque` (linked blocks) |
| Hash table             | array indexed by `hash(key)`                        | environment                        | `dict`, `set`       |
:::

::: {.aside}
See [R Internals](https://cran.r-project.org/doc/manuals/r-release/R-ints.html) and [CPython Internals](https://github.com/zpoint/CPython-Internals) for the gory details.
:::



## Vectors

Materialized numeric vectors in R and contiguous numeric NumPy arrays store fixed-size values together, giving O(1) element access. NumPy views can have gaps or reversed strides.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
object.size(numeric(1e6))
object.size(integer(1e6))
```
:::

::: {.column width='50%'}
```{python}
np.zeros(1_000_000).nbytes
np.zeros(1_000_000, dtype="int64").nbytes
np.zeros(1_000_000, dtype="int32").nbytes
```
:::
::::

::: {.aside}
R doubles use 8 bytes per element; integers use 4. `object.size()` includes object overhead; NumPy's `.nbytes` counts only element storage.
:::


## Growing a vector

Repeatedly appending with `c(x, i)` copies the existing elements each time, giving O($n^2$) work overall.


::: {.xsmall}
```{r}
n = 5e4
x = c()
system.time(
  for (i in seq_len(n)) x = c(x, i)
)
```
:::

Indexed growth can reuse spare capacity, but still needs occasional copying. Preallocation avoids repeated growth.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
x = numeric()
system.time(
  for (i in seq_len(n)) x[i] = i
)
```
:::

::: {.column width='50%' .fragment}
```{r}
x = numeric(n)
system.time(
  for (i in seq_len(n)) x[i] = i
)
```
:::
::::



## Array (Vector)

<iframe data-src="https://visualgo.net/en/array" width="100%" height="450px" style="border:1px solid;border-radius: 5px;" data-external="1">
</iframe>

::: {.aside}
<https://visualgo.net/en/array> - Materialized numeric R vectors and contiguous numeric NumPy arrays store same-typed values together. A Python `list` is a contiguous array of *pointers* to objects, which is why it can be heterogeneous and why it uses far more memory.
:::

## Generic vectors

R and Python lists store references to objects. Elements can have different types and sizes, while indexing remains O(1). References and object headers add memory overhead.

* Numeric vector - values stored contiguously:

  ::: {.small}
  ```
  [ 1.0 | 2.0 | 3.0 ]
  ```
  :::

* List -  *references* stored contiguously

  ::: {.small}
  ```
  [  P  |  P  |  P  ]
     ↓     ↓     ↓
     1   "abc"  TRUE
  ```
  :::


## deques

A `deque` (double-ended queue) supports O(1) additions and removals at either end. CPython stores it as a doubly linked chain of fixed-size blocks. Indexing is O(1) near either end and O(n) in the middle.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
from collections import deque
x = deque(range(3))
x.appendleft(-1)
x
```
```{python}
x.popleft()
x
```
:::

::: {.column width='50%' .fragment}
```{python}
x = deque(range(3), maxlen=4)
x.append(10)
x.append(11)
x
```
```{python}
x.appendleft(-1)
x
```
:::
::::

::: {.aside}
With `maxlen` set, adding to a full deque drops an element from the opposite end - a rolling window of the last $n$ items. `list.pop(0)` and `list.insert(0, v)` have to shift every remaining element.
:::


## Linked list

<iframe data-src="https://visualgo.net/en/list" width="100%" height="450px" style="border:1px solid;border-radius: 5px;" data-external="1">
</iframe>

::: {.aside}
<https://visualgo.net/en/list> - CPython's `deque` links blocks of elements. R uses *pairlists* (linked lists) internally for function arguments and attributes.
:::


## Hashing

A hash table is an array of *buckets* - a key's `hash()` picks its bucket, so lookup, insertion and deletion are O(1) on average with no scanning. Two keys landing in the same bucket is a *collision*, and the table is resized as it fills to keep collisions rare.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{python}
hash(1), hash(1.0), hash(True)
hash(2.5)
hash((1, 2))
```
```{python}
keys = (1, 9, 17, 2.5)
[hash(k) % 8 for k in keys]
```
:::

::: {.column width='50%' .fragment}
```{r}
h = new.env(size = 8L)
for (k in letters[1:5]) h[[k]] = k
str(env.profile(h))
```
```{r}
for (k in letters[6:20]) h[[k]] = k
env.profile(h)$size
```
:::
::::

::: {.aside}
Equal keys must hash equally, which is why `1`, `1.0`, and `True` are the same `dict` key. String hashes are randomized per process, so `hash("abc")` (and the iteration order of a set of strings) changes between runs. `env.profile()` reports the bucket sizes of an environment's table - R chains colliding keys within a bucket, CPython probes for another free slot.
:::


## Hash table

<iframe data-src="https://visualgo.net/en/hashtable" width="100%" height="450px" style="border:1px solid;border-radius: 5px;" data-external="1">
</iframe>

::: {.aside}
<https://visualgo.net/en/hashtable> - Python's `dict` and `set` are hash tables, as are R environments. R's `match()`, `%in%`, `unique()`, and `duplicated()` build a hash table of one argument, which is why they are fast.
:::


## Operation costs

::: {.small}
| Operation                 | `list`                  | `deque`                   | `dict`                | `set`                 |
|:--------------------------|:------------------------|:--------------------------|:----------------------|:----------------------|
| Copy                      | `.copy()`: O(n)         | `.copy()`: O(n)           | `.copy()`: O(n)       | `.copy()`: O(n)       |
| Get / set by index or key | `s[i]`: O(1)            | `s[i]`: O(n)              | `d[k]`: O(1)          | —                     |
| Append / add              | `.append(v)`: O(1)\*    | `.append(v)`: O(1)        | `d[k] = v`: O(1)\*    | `.add(v)`: O(1)\*     |
| Pop from end / remove     | `.pop()`: O(1)\*        | `.pop()`: O(1)            | `.pop(k)`: O(1)       | `.remove(v)`: O(1)    |
| Insert at front           | `.insert(0, v)`: O(n)   | `.appendleft(v)`: O(1)    | —                     | —                     |
| Remove from front         | `.pop(0)`: O(n)         | `.popleft()`: O(1)        | —                     | —                     |
| Insert / delete in middle | O(n)                    | O(n)                      | —                     | —                     |
| Membership                | `v in s`: O(n)          | `v in s`: O(n)            | `k in d`: O(1)        | `v in s`: O(1)        |
:::

::: {.aside}
CPython costs for n elements. 

\*Amortized - list append and end pop, and hash table insertion, occasionally trigger a resize. Deque indexing is O(1) near either end. Hash table costs are averages that assume constant-cost hashing and equality, an individual operation can be O(n). Dictionary membership tests keys. See <https://wiki.python.org/moin/TimeComplexity>.
:::


## Exercise 2

For each of the following, suggest a data structure in Python and R. Explain which operations your choice makes efficient, and state any assumptions.

::: {.small}
* A collection of 100 integers used repeatedly in elementwise arithmetic.

* A queue (first in, first out) of customer records.

* A stack (first in, last out) of customer records.

* A count of word occurrences within a document.

* The heights of the bars in a histogram with even bin widths.

* Checking whether each of a million words is in a list of 50 stop words.
:::

```{r}
#| echo: false
countdown::countdown(minutes = 3)
```


## Measuring vector growth

Reasoning about complexity tells you what to expect, measuring confirms it

::: {.panel-tabset}

### Code


::: {.columns .xsmall}

::: {.column}
```{r}
grow_concat = function(n) {
  x = c()
  for (i in seq_len(n)) x = c(x, i)
}
grow_index = function(n) {
  x = numeric()
  for (i in seq_len(n)) x[i] = i
}
grow_prealloc = function(n) {
  x = numeric(n)
  for (i in seq_len(n)) x[i] = i
}
```
:::

::: {.column}
```{r}
#| message: false
#| warning: false
#| cache: true

res = bench::press(
  n = c(1e3, 2.5e3, 5e3, 1e4, 2.5e4, 5e4),
  bench::mark(
    concat   = grow_concat(n),
    index    = grow_index(n),
    prealloc = grow_prealloc(n),
    check = FALSE
  )
)
```
:::
:::



::: {.xsmall}

:::

### Linear

```{r}
#| echo: false
#| fig-width: 9
#| fig-height: 4.25
#| fig-align: center
timings = res |>
  dplyr::mutate(method = as.character(expression), median = as.numeric(median))

g = ggplot2::ggplot(timings, ggplot2::aes(x = n, y = median, color = method)) +
  ggplot2::geom_line() +
  ggplot2::geom_point() +
  ggplot2::labs(x = "n", y = "median time (s)", color = NULL) +
  ggplot2::theme_minimal(base_size = 16)
g
```

### Log

```{r}
#| echo: false
#| fig-width: 9
#| fig-height: 4.25
#| fig-align: center
g +
  ggplot2::scale_x_log10() +
  ggplot2::scale_y_log10(labels = scales::label_log())
```

:::

::: {.aside}
`bench::mark()` repeats each expression and reports the distribution of timings, `bench::press()` runs it over a grid of parameters.
:::


# Comparing R & Python {visibility="uncounted"}

## Container summary {visibility="uncounted"}

::: {.small}
| Python                     | Properties                              | R                                           |
|:---------------------------|:----------------------------------------|:--------------------------------------------|
| `list`                     | ordered, mutable, heterogeneous         | `list()`                                    |
| `tuple`                    | ordered, immutable, heterogeneous       | `list()` or atomic vector (copy on modify)  |
| `dict`                     | key lookup, mutable, insertion ordered  | named `list()`, `new.env()` for hashing     |
| `set`                      | unique, unordered, mutable              | vector + `unique()`, `union()`, `%in%`, ... |
| `deque`                    | fast at both ends                       | none (use a list or vector)                 |
| `range`                    | lazy integer sequence, immutable        | `seq_len()`, `:` (often stored compactly via ALTREP) |
| NumPy `ndarray`            | homogeneous, vectorized                 | atomic vector, matrix, array                |
:::


## Takeaways {visibility="uncounted"}

* **Tuples** are immutable sequences—use them for fixed groups of values and unpacking.

* **Dictionaries** map hashable keys to values, with O(1) lookup on average. R offers named lists and hash-based environments.

* **Sets** provide O(1) membership on average and set algebra. In R, use `%in%`, `union()`, `intersect()`, and `setdiff()`.

* **Shallow copies** share nested objects. Use `deepcopy()` when nested objects need independent mutation.

* **Mutable defaults** are shared between calls. Use `None` and create the object inside the function.

* Choose containers for their operations: deques for queues, sets for membership, and preallocated R vectors when the size is known. Measure when performance matters.
