---
title: "Lists, S3, &<br/>Python Classes"
subtitle: "Lecture 05"
author: "Dr. Colin Rundel"
footer: "Sta 523 - Fall 2026"
format:
  revealjs:
    theme: slides.scss
    transition: fade
    slide-number: true
    self-contained: true
execute:
  echo: true
  warning: true
engine: knitr
---


```{r setup}
#| message: false
#| warning: false
#| include: false
options(
  width = 80
)
```


# R Lists

## Lists

Lists are R's other vector type (generic). Unlike atomic vectors they are *heterogeneous* - each element can be any R object: atomic vectors, other lists, functions, etc. 


:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
list("A", (1:4)/2, list(1L), sum)
```
:::

::: {.column width='50%'}
```{python}
["A", [0.5, 1.0, 1.5, 2.0], [1], sum]
```
:::
::::

::: {.aside}
`is.vector()` returns `TRUE` for both atomic and generic vectors, instead use `is.atomic()` and `is.list()`.
:::


## List structure

The printed form of a list is verbose. `str()` gives a compact summary of any R object's structure and is particularly useful for lists.

:::: {.columns .small}
::: {.column width='50%'}
```{r}
str(c(1, 2))
str(1:100)
str("A")
```
:::

::: {.column width='50%' .fragment}
```{r}
str( list(
  "A", c(TRUE, FALSE),
  (1:4)/2, list(TRUE, 1),
  function(x) x^2
) )
```
:::
::::


## Nested lists

Lists can contain other lists, so they do not have to be flat. 

This makes them a natural way of representing tree-like data (e.g. JSON).

:::: {.columns .small}
::: {.column width='50%'}
```{r}
x = list(1, list(2, list(3, 4), 5))
str(x)
```
:::

::: {.column width='50%'}
```{python}
x = [1, [2, [3, 4], 5]]
x
```
:::
::::

::: {.aside}
Lists are not the choice for JSON in Python because their elements cannot be named
:::


## Named lists

Elements of a list (or atomic vector) can be named. Names help avoid magic numbers when accessing elements (more readable code).

A valid name starts with a letter or `.` (not followed by a digit) and contains only `[A-Za-z0-9._]`.

::: {.xsmall}
```{r}
x = list(A = 1, B = list(C = 2, D = 3))
str(x)
names(x)
```
:::

. . .

Any other name must be surrounded with backticks,

::: {.xsmall}
```{r}
list("knock knock" = "who's there?")
```
:::


## `[` vs `[[`

R has two addition subsetting operators. 

* `[[` extracts a *single element*

* `$` is shorthand for named lookup with `[[`

::: {.xsmall}
```{r}
y = list(a = 1, b = 4, c = 7:9)
```
:::

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
y[2]
str( y[2] )
y["b"]
str( y["b"] )
```
:::

::: {.column width='50%' .fragment}
```{r}
y[[2]]
str( y[[2]] )
y[["b"]]
str( y[["b"]] )
y$b
```
:::
::::




## Hadley's analogy

```{r}
#| echo: false
#| fig-align: center
#| out-width: 37%
knitr::include_graphics("imgs/list_train1.png")
```
. . .
```{r}
#| echo: false
#| fig-align: center
#| out-width: 37%
knitr::include_graphics("imgs/list_train2.png")
```
. . .
```{r}
#| echo: false
#| fig-align: center
#| out-width: 37%
knitr::include_graphics("imgs/list_train3.png")
```

::: {.aside}
From Advanced R - [Chapter 4.3](https://adv-r.hadley.nz/subsetting.html#subset-single)
:::


## `[[` details

`[[` selects by position or by name and only ever returns one element. 

::: {.small}
```{r}
y = list(a = 1, b = 4, c = 7:9)
```
:::

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
y[[3]]
y[["c"]]
```
:::

::: {.column width='50%' .fragment}
```{r}
#| error: true
y[[4]]
y[["d"]]
```
:::
::::


Vectors of length > 1 are interpreted as *recursive* indexing - `y[[c(3, 2)]]` is `y[[3]][[2]]`.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
y[[c(3, 2)]]
```
:::

::: {.column width='50%' .fragment}
```{r}
#| error: true
y[[1:2]]
```
:::
::::


## `$` subsetting

`$` is a shorthand for `[[` with a *literal* name, so `y$c` is equivalent to `y[["c"]]`. It only works with lists and objects built using lists (e.g.data frames) and it *partially matches names*.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
y = list(abc = 1, def = 5)
```
```{r}
#| error: true
y[["abc"]]
y$abc
y$a
y[["a"]]
```
:::

::: {.column width='50%' .fragment}

```{r}
x = c(abc = 1, def = 5)
```
```{r}
#| error: true
x[["abc"]]
x$abc
```
:::
::::

. . .

A common error is using `$` with a variable that holds a name,

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
name = "def"
y[[name]]
```
:::

::: {.column width='50%'}
```{r}
y$name
```
:::
::::


## Modifying lists

Assignment with `[[` or `$` replaces an element or adds a new one. 

Assigning `NULL` removes an element.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
x = list(a = 1, b = 2)
x$c = "new"
x[["a"]] = 100
str(x)
```
```{r}
x$b = NULL
str(x)
```
:::

::: {.column .fragment width='50%'}
```{python}
x = [1, 2]
x.append("new")
x[0] = 100
x
```
```{python}
del x[1]
x
```
:::
::::


## Lists and atomic vectors

Combining an atomic vector with a list via `c()` produces a list (the more generic type). 

`unlist()` goes the other way, flattening a list into an atomic vector with the usual coercion rules.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
str( c(1, list(4, list(6, 7))) )
```
:::

::: {.column width='50%' .fragment}
```{r}
unlist( list(1:3, list(4:5, 6)) )
unlist( list(1, list(2, list(3, "Hello"))) )
```
:::
::::


## Exercise 1

Represent the following JSON data as a list in R (we will revisit this in Python next week).

:::: {.columns}
::: {.column width='60%'}
::: {.xsmall}
```json
{
  "firstName": "John",
  "lastName": "Smith",
  "age": 25,
  "address":
  {
    "streetAddress": "21 2nd Street",
    "city": "New York",
    "state": "NY",
    "postalCode": 10021
  },
  "phoneNumber":
  [ {
      "type": "home",
      "number": "212 555-1239"
    },
    {
      "type": "fax",
      "number": "646 555-4567"
  } ]
}
```
:::
:::

::: {.column width='40%'}
::: {.small}
Once you have the list:

* Extract the fax number using `$` and `[[`.

* Add a `"mobile"` phone number.

* Remove the `age` element.
:::
:::
::::


# Attributes

## Attributes

Attributes are metadata attached to an R object. Some attributes are special (e.g. `names`, `dim`, `dimnames`, `class`, `levels`, etc.) because they modoify how the object behaves / is treated.

. . .

Attributes are stored as a *named list* attached to the object, accessed via `attributes()` or `attr()`.

::: {.small}
```{r}
(x = c(L = 1, M = 2, N = 3))
```
:::

:::: {.columns .small}
::: {.column width='50%'}
```{r}
str( attributes(x) )
```
:::

::: {.column width='50%' .fragment}
```{r}
attr(x, "names")
attr(x, "other")
```
:::
::::


## Setting attributes

Most important attributes have helper functions for getting and setting (`names()`, `dim()`, `class()`, `levels()`),

:::: {.columns .small}
::: {.column width='50%'}
```{r}
names(x) = c("Z", "Y", "X")
x
names(x) = 1:3
x
attributes(x)
```
:::

::: {.column width='50%' .fragment}
```{r}
attr(x, "other") = "anything"
x
str( attributes(x) )
```
:::
::::

::: {.aside}
Note that `names(x) = 1:3` silently coerced the integers to character, since the `names` attribute must be a character vector of the same length as the object.
:::


## Factors

Factors are how R represents categorical data - a variable with a discrete set of possible values (called levels).

::: {.small}
```{r}
(x = factor(c("Sunny", "Cloudy", "Rainy", "Cloudy", "Cloudy")))
```
:::

. . .

::: {.small}
```{r}
str(x)
```
:::

. . .

:::: {.columns .small}
::: {.column width='50%'}
```{r}
typeof(x)
mode(x)
```
:::

::: {.column width='50%'}
```{r}
class(x)
levels(x)
```
:::
::::


## Composition

A factor is just an integer vector with two attributes: `levels` and `class`.

::: {.small}
```{r}
str( attributes(x) )
unclass(x)
```
:::

. . .

We can build our own from scratch using `attr()`,

::: {.small}
```{r}
y = c(3L, 1L, 2L, 1L, 1L)
attr(y, "levels") = c("Cloudy", "Rainy", "Sunny")
attr(y, "class") = "factor"
y
```
:::


## Building objects with `structure()`

Setting attributes one at a time is clunky - `structure()` attaches any number of attributes to an object in a single call and is the standard way to construct objects like this.

::: {.small}
```{r}
( y = structure(
    c(3L, 1L, 2L, 1L, 1L),
    levels = c("Cloudy", "Rainy", "Sunny"),
    class = "factor"
) )
```
:::

. . .

::: {.small}
```{r}
class(y)
is.factor(y)
identical(x, y)
```
:::


## Factors are integer vectors?

Knowing that factors are stored as integers explains some of their more surprising behaviors,

:::: {.columns .small}
::: {.column width='50%'}
```{r}
#| error: true
x + 1
is.integer(x)
is.numeric(x)
```
:::

::: {.column width='50%' .fragment}
```{r}
#| error: true
as.integer(x)
as.character(x)
as.logical(x)
```
:::
::::

. . .

::: {.columns .small}
::: {.column}
```{r}
as.numeric(factor(c("10", "20")))
```
:::

::: {.column}
```{r}
as.numeric(
  as.character(factor(c("10", "20")))
)
```
:::
:::


::: {.aside}
Most of these are pretty sensible choices - this was not always the case
:::


# S3 Object System

## `class`

The `class` attribute adds a layer on top of R's type hierarchy - previously we saw `typeof()` and `mode()`.

```{r}
#| echo: false
f = function(x) x^2
x = factor("A")
l = list(1, "A")
m = matrix(1:4, 2)
```

::: {.small}
 value             | `typeof()`       | `mode()`       | `class()`
:------------------|:-----------------|:---------------|:---------------
`TRUE`             | `r typeof(TRUE)` | `r mode(TRUE)` | `r class(TRUE)`
`1`                | `r typeof(1)`    | `r mode(1)`    | `r class(1)`
`1L`               | `r typeof(1L)`   | `r mode(1L)`   | `r class(1L)`
`"A"`              | `r typeof("A")`  | `r mode("A")`  | `r class("A")`
`NULL`             | `r typeof(NULL)` | `r mode(NULL)` | `r class(NULL)`
`list(1, "A")`     | `r typeof(l)`    | `r mode(l)`    | `r class(l)`
`factor("A")`      | `r typeof(x)`    | `r mode(x)`    | `r class(x)`
`matrix(1:4, 2)`   | `r typeof(m)`    | `r mode(m)`    | `r paste(class(m), collapse = ", ")`
`function(x) x^2`  | `r typeof(f)`    | `r mode(f)`    | `r class(f)`
`sum`              | `r typeof(sum)`  | `r mode(sum)`  | `r class(sum)`
:::

::: {.aside}
Without a `class` attribute an object has an *implicit* class derived from its type - for most vectors this matches `mode()`.
:::


## S3 class specialization

::: {.small}
```{r}
x = c("A", "B", "A", "C")
```
:::

. . .

::: {.small}
```{r}
print( x )
```
:::

. . .

::: {.small}
```{r}
print( factor(x) )
```
:::

. . .

::: {.small}
```{r}
print( unclass( factor(x) ) )
```
:::

. . .

::: {.small}
```{r}
print.default( factor(x) )
```
:::


## What's up with `print`? {.scrollable}

::: {.xsmall}
```{r}
print
```
:::

. . .

::: {.medium}
`print` does no printing itself - `UseMethod()` looks at the class of `x` and calls the matching *method*. For a factor that is `print.factor()`, for everything without a more specific method `print.default()` is used.
:::

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
print.default
```
:::

::: {.column width='50%'}
```{r}
print.factor
```
:::
::::


## Generics are everywhere

Many of the base R functions you use regularly are S3 generics whose only job is to dispatch on class,

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
mean
summary
```
:::

::: {.column width='50%'}
```{r}
plot
t.test
```
:::
::::

. . .

Other generics, such as `sum`, dispatch without an explicit call to `UseMethod()`:

::: {.xsmall}
```{r}
sum
```
:::

::: {.aside}
`sum` is a primitive implemented in C. It still performs S3 dispatch internally, rather than calling `UseMethod()` in its visible R definition.
:::


## What is S3?

<br/>

> S3 is R’s first and simplest OO system. S3 is R's first and simplest OO system. S3 is informal and ad hoc, but there is a certain elegance in its minimalism: you can't take away any part of it and still have a useful OO system.
>
> - Hadley Wickham, Advanced R

::: {.aside}
S3 is the only OO system used in the base and stats packages and the most common on CRAN. It should not be confused with R's other object-oriented systems: S4, Reference classes, R6, and S7.
:::


## Dispatch {.scrollable}

S3 dispatch is very simple - a generic calls `UseMethod("func")`, which searches for a function named `<func>.<class>`, trying each class of the first argument in order and falling back to `<func>.default` if none is found.

. . .

`methods()` lists the methods available for a generic, or all the methods available for a class,

::: {.xsmall}
```{r}
methods("summary")
```
:::

. . .

::: {.xsmall}
```{r}
methods(class = "factor")
```
:::


## Adding methods

Because dispatch is by *name*, adding a method for a new class just requires defining a new function with the proper name - no registration is needed.

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
( x = structure(
    c(1, 2, 3),
    class = "class_A") )
```
:::

::: {.column width='50%'}
```{r}
( y = structure(
    c(6, 5, 4),
    class = "class_B") )
```
:::
::::

. . .

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
print.class_A = function(x, ...) {
  cat("(Class A) ")
  print.default(unclass(x))
}
```
:::

::: {.column width='50%'}
```{r}
print.class_B = function(x, ...) {
  cat("(Class B) ")
  print.default(unclass(x))
}
```
:::
::::

. . .

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
print(x)
```
:::

::: {.column width='50%'}
```{r}
print(y)
```
:::
::::


. . .

:::: {.columns .xsmall}
::: {.column width='50%'}
```{r}
class(x) = "class_B"
print(x)
```
:::

::: {.column width='50%'}
```{r}
class(y) = "class_A"
print(y)
```
:::
::::


## Writing a good `print` method

In general, methods should match the generic's signature (`x, ...`). For print you should then manage output via `cat()` or `print()`, and return the input *invisibly* so that `print(x)` can be used in a pipeline.

::: {.small}
```{r}
pct = structure(c(0.12, 0.5, 1), class = "percent")

print.percent = function(x, ...) {
  print(paste0(unclass(x) * 100, "%"), quote = FALSE)
  invisible(x)
}
```
:::

. . .

:::: {.columns .small}
::: {.column width='50%'}
```{r}
pct
z = print(pct)
```
:::

::: {.column width='50%' .fragment}
```{r}
unclass(z)
```
:::
::::

::: {.aside}
This goes for any new method - match the signature and conform to the common expectations of the other method implementations.
:::

## Defining a new generic

A generic is just a function that calls `UseMethod()`. By convention the first argument is the object dispatched on and `...` is included so that methods can add arguments.

::: {.small}
```{r}
shuffle = function(x, ...) {
  UseMethod("shuffle")
}
```
:::

. . .

::: {.small}
```{r}
shuffle.default = function(x, ...) {
  stop("Class ", class(x), " is not supported by shuffle.", call. = FALSE)
}
```
:::

. . .

::: {.small}
```{r}
shuffle.factor = function(x, ...) {
  factor( sample(as.character(x)), levels = sample(levels(x)) )
}
```
```{r}
shuffle.integer = function(x, ...) {
  sample(x)
}
```
:::


## Shuffle results

::: {.small}
```{r}
shuffle( 1:10 )
```
:::

. . .

::: {.small}
```{r}
shuffle( factor(c("A", "B", "C", "A")) )
```
:::

. . .

::: {.small}
```{r}
#| error: true
shuffle( c(1, 2, 3, 4, 5) )
```
:::

. . .

::: {.small}
```{r}
#| error: true
shuffle( letters[1:5] )
```
:::

. . .

::: {.small}
```{r}
methods("shuffle")
```
:::


## Implicit classes

Objects without a `class` attribute dispatch on an *implicit* class vector that is more detailed than what `class()` reports.

::: {.xsmall}
```{r}
report = function(x) {
  UseMethod("report")
}
report.default = function(x) paste0("Class ", class(x), " does not have a method defined.")
report.integer = function(x) "I'm an integer!"
report.double  = function(x) "I'm a double!"
report.numeric = function(x) "I'm a numeric!"
```
:::

. . .

:::: {.columns .xsmall}
::: {.column width='33%'}
```{r}
report(1)
report(1L)
report("1")
```
:::

::: {.column width='33%' .fragment}
```{r}
rm(report.integer, report.double)
report(1)
report(1L)
```
:::

::: {.column width='33%' .fragment}
```{r}
rm(report.numeric)
report(1)
report(1L)
```
:::
::::


## Why?

::: {.medium}
From `UseMethod`'s R documentation:

> If the object does not have a class attribute, it has an implicit class. Matrices and arrays have class "matrix" or "array" followed by the class of the underlying vector. Most vectors have class the result of `mode(x)`, except that integer vectors have class `c("integer", "numeric")` and real vectors have class `c("double", "numeric")`.
:::

. . .

::: {.medium}
The implicit class vector can be inspected with `.class2()`,
:::

::: {.xsmall}
```{r}
.class2(1)
.class2(1L)
.class2(matrix(1:4, 2))
```
:::

::: {.aside}
`report(1)` tries `report.double`, then `report.numeric`, then `report.default` - which is why a method for `double` is found even though `class(1)` says `"numeric"`.
:::


