Lists, S3, &
Python Classes

Lecture 05

Dr. Colin Rundel

R Lists

Lists

Lists are R’s other vector type (generic). Unlike atomic vectors they are heterogeneous - each element can be any R object: atomic vectors, other lists, functions, etc.

list("A", (1:4)/2, list(1L), sum)
[[1]]
[1] "A"

[[2]]
[1] 0.5 1.0 1.5 2.0

[[3]]
[[3]][[1]]
[1] 1


[[4]]
function (..., na.rm = FALSE)  .Primitive("sum")
["A", [0.5, 1.0, 1.5, 2.0], [1], sum]
['A', [0.5, 1.0, 1.5, 2.0], [1], <built-in function sum>]

List structure

The printed form of a list is verbose. str() gives a compact summary of any R object’s structure and is particularly useful for lists.

str(c(1, 2))
 num [1:2] 1 2
str(1:100)
 int [1:100] 1 2 3 4 5 6 7 8 9 10 ...
str("A")
 chr "A"
str( list(
  "A", c(TRUE, FALSE),
  (1:4)/2, list(TRUE, 1),
  function(x) x^2
) )
List of 5
 $ : chr "A"
 $ : logi [1:2] TRUE FALSE
 $ : num [1:4] 0.5 1 1.5 2
 $ :List of 2
  ..$ : logi TRUE
  ..$ : num 1
 $ :function (x)  

Nested lists

Lists can contain other lists, so they do not have to be flat.

This makes them a natural way of representing tree-like data (e.g. JSON).

x = list(1, list(2, list(3, 4), 5))
str(x)
List of 2
 $ : num 1
 $ :List of 3
  ..$ : num 2
  ..$ :List of 2
  .. ..$ : num 3
  .. ..$ : num 4
  ..$ : num 5
x = [1, [2, [3, 4], 5]]
x
[1, [2, [3, 4], 5]]

Named lists

Elements of a list (or atomic vector) can be named. Names help avoid magic numbers when accessing elements (more readable code).

A valid name starts with a letter or . (not followed by a digit) and contains only [A-Za-z0-9._].

x = list(A = 1, B = list(C = 2, D = 3))
str(x)
List of 2
 $ A: num 1
 $ B:List of 2
  ..$ C: num 2
  ..$ D: num 3
names(x)
[1] "A" "B"

Any other name must be surrounded with backticks,

list("knock knock" = "who's there?")
$`knock knock`
[1] "who's there?"

[ vs [[

R has two addition subsetting operators.

  • [[ extracts a single element

  • $ is shorthand for named lookup with [[

y = list(a = 1, b = 4, c = 7:9)
y[2]
$b
[1] 4
str( y[2] )
List of 1
 $ b: num 4
y["b"]
$b
[1] 4
str( y["b"] )
List of 1
 $ b: num 4
y[[2]]
[1] 4
str( y[[2]] )
 num 4
y[["b"]]
[1] 4
str( y[["b"]] )
 num 4
y$b
[1] 4

Hadley’s analogy

[[ details

[[ selects by position or by name and only ever returns one element.

y = list(a = 1, b = 4, c = 7:9)
y[[3]]
[1] 7 8 9
y[["c"]]
[1] 7 8 9
y[[4]]
Error in `y[[4]]`:
! subscript out of bounds
y[["d"]]
NULL

Vectors of length > 1 are interpreted as recursive indexing - y[[c(3, 2)]] is y[[3]][[2]].

y[[c(3, 2)]]
[1] 8
y[[1:2]]
Error in `y[[1:2]]`:
! subscript out of bounds

$ subsetting

$ is a shorthand for [[ with a literal name, so y$c is equivalent to y[["c"]]. It only works with lists and objects built using lists (e.g.data frames) and it partially matches names.

y = list(abc = 1, def = 5)
y[["abc"]]
[1] 1
y$abc
[1] 1
y$a
[1] 1
y[["a"]]
NULL
x = c(abc = 1, def = 5)
x[["abc"]]
[1] 1
x$abc
Error in `x$abc`:
! $ operator is invalid for atomic vectors

A common error is using $ with a variable that holds a name,

name = "def"
y[[name]]
[1] 5
y$name
NULL

Modifying lists

Assignment with [[ or $ replaces an element or adds a new one.

Assigning NULL removes an element.

x = list(a = 1, b = 2)
x$c = "new"
x[["a"]] = 100
str(x)
List of 3
 $ a: num 100
 $ b: num 2
 $ c: chr "new"
x$b = NULL
str(x)
List of 2
 $ a: num 100
 $ c: chr "new"
x = [1, 2]
x.append("new")
x[0] = 100
x
[100, 2, 'new']
del x[1]
x
[100, 'new']

Lists and atomic vectors

Combining an atomic vector with a list via c() produces a list (the more generic type).

unlist() goes the other way, flattening a list into an atomic vector with the usual coercion rules.

str( c(1, list(4, list(6, 7))) )
List of 3
 $ : num 1
 $ : num 4
 $ :List of 2
  ..$ : num 6
  ..$ : num 7
unlist( list(1:3, list(4:5, 6)) )
[1] 1 2 3 4 5 6
unlist( list(1, list(2, list(3, "Hello"))) )
[1] "1"     "2"     "3"     "Hello"

Exercise 1

Represent the following JSON data as a list in R (we will revisit this in Python next week).

{
  "firstName": "John",
  "lastName": "Smith",
  "age": 25,
  "address":
  {
    "streetAddress": "21 2nd Street",
    "city": "New York",
    "state": "NY",
    "postalCode": 10021
  },
  "phoneNumber":
  [ {
      "type": "home",
      "number": "212 555-1239"
    },
    {
      "type": "fax",
      "number": "646 555-4567"
  } ]
}

Once you have the list:

  • Extract the fax number using $ and [[.

  • Add a "mobile" phone number.

  • Remove the age element.

Attributes

Attributes

Attributes are metadata attached to an R object. Some attributes are special (e.g. names, dim, dimnames, class, levels, etc.) because they modoify how the object behaves / is treated.

Attributes are stored as a named list attached to the object, accessed via attributes() or attr().

(x = c(L = 1, M = 2, N = 3))
L M N 
1 2 3 
str( attributes(x) )
List of 1
 $ names: chr [1:3] "L" "M" "N"
attr(x, "names")
[1] "L" "M" "N"
attr(x, "other")
NULL

Setting attributes

Most important attributes have helper functions for getting and setting (names(), dim(), class(), levels()),

names(x) = c("Z", "Y", "X")
x
Z Y X 
1 2 3 
names(x) = 1:3
x
1 2 3 
1 2 3 
attributes(x)
$names
[1] "1" "2" "3"
attr(x, "other") = "anything"
x
1 2 3 
1 2 3 
attr(,"other")
[1] "anything"
str( attributes(x) )
List of 2
 $ names: chr [1:3] "1" "2" "3"
 $ other: chr "anything"

Factors

Factors are how R represents categorical data - a variable with a discrete set of possible values (called levels).

(x = factor(c("Sunny", "Cloudy", "Rainy", "Cloudy", "Cloudy")))
[1] Sunny  Cloudy Rainy  Cloudy Cloudy
Levels: Cloudy Rainy Sunny
str(x)
 Factor w/ 3 levels "Cloudy","Rainy",..: 3 1 2 1 1
typeof(x)
[1] "integer"
mode(x)
[1] "numeric"
class(x)
[1] "factor"
levels(x)
[1] "Cloudy" "Rainy"  "Sunny" 

Composition

A factor is just an integer vector with two attributes: levels and class.

str( attributes(x) )
List of 2
 $ levels: chr [1:3] "Cloudy" "Rainy" "Sunny"
 $ class : chr "factor"
unclass(x)
[1] 3 1 2 1 1
attr(,"levels")
[1] "Cloudy" "Rainy"  "Sunny" 

We can build our own from scratch using attr(),

y = c(3L, 1L, 2L, 1L, 1L)
attr(y, "levels") = c("Cloudy", "Rainy", "Sunny")
attr(y, "class") = "factor"
y
[1] Sunny  Cloudy Rainy  Cloudy Cloudy
Levels: Cloudy Rainy Sunny

Building objects with structure()

Setting attributes one at a time is clunky - structure() attaches any number of attributes to an object in a single call and is the standard way to construct objects like this.

( y = structure(
    c(3L, 1L, 2L, 1L, 1L),
    levels = c("Cloudy", "Rainy", "Sunny"),
    class = "factor"
) )
[1] Sunny  Cloudy Rainy  Cloudy Cloudy
Levels: Cloudy Rainy Sunny
class(y)
[1] "factor"
is.factor(y)
[1] TRUE
identical(x, y)
[1] TRUE

Factors are integer vectors?

Knowing that factors are stored as integers explains some of their more surprising behaviors,

x + 1
Warning in Ops.factor(x, 1): '+' not meaningful for factors
[1] NA NA NA NA NA
is.integer(x)
[1] FALSE
is.numeric(x)
[1] FALSE
as.integer(x)
[1] 3 1 2 1 1
as.character(x)
[1] "Sunny"  "Cloudy" "Rainy"  "Cloudy" "Cloudy"
as.logical(x)
[1] NA NA NA NA NA
as.numeric(factor(c("10", "20")))
[1] 1 2
as.numeric(
  as.character(factor(c("10", "20")))
)
[1] 10 20

S3 Object System

class

The class attribute adds a layer on top of R’s type hierarchy - previously we saw typeof() and mode().

value typeof() mode() class()
TRUE logical logical logical
1 double numeric numeric
1L integer numeric integer
"A" character character character
NULL NULL NULL NULL
list(1, "A") list list list
factor("A") integer numeric factor
matrix(1:4, 2) integer numeric matrix, array
function(x) x^2 closure function function
sum builtin function function

S3 class specialization

x = c("A", "B", "A", "C")
print( x )
[1] "A" "B" "A" "C"
print( factor(x) )
[1] A B A C
Levels: A B C
print( unclass( factor(x) ) )
[1] 1 2 1 3
attr(,"levels")
[1] "A" "B" "C"
print.default( factor(x) )
[1] 1 2 1 3

What’s up with print?

print
function (x, ...) 
UseMethod("print")
<bytecode: 0x75fffdccb8>
<environment: namespace:base>

print does no printing itself - UseMethod() looks at the class of x and calls the matching method. For a factor that is print.factor(), for everything without a more specific method print.default() is used.

print.default
function (x, digits = NULL, quote = TRUE, na.print = NULL, print.gap = NULL, 
    right = FALSE, max = NULL, width = NULL, useSource = TRUE, 
    ...) 
{
    args <- pairlist(digits = digits, quote = quote, na.print = na.print, 
        print.gap = print.gap, right = right, max = max, width = width, 
        useSource = useSource, ...)
    missings <- c(missing(digits), missing(quote), missing(na.print), 
        missing(print.gap), missing(right), missing(max), missing(width), 
        missing(useSource))
    .Internal(print.default(x, args, missings))
}
<bytecode: 0x75fcd11968>
<environment: namespace:base>
print.factor
function (x, quote = FALSE, max.levels = NULL, width = getOption("width"), 
    ...) 
{
    ord <- is.ordered(x)
    if (length(x) == 0L) 
        cat(if (ord) 
            "ordered"
        else "factor", "()\n", sep = "")
    else {
        xx <- character(length(x))
        xx[] <- as.character(x)
        keepAttrs <- setdiff(names(attributes(x)), c("levels", 
            "class"))
        attributes(xx)[keepAttrs] <- attributes(x)[keepAttrs]
        print(xx, quote = quote, ...)
    }
    maxl <- max.levels %||% TRUE
    if (maxl) {
        n <- length(lev <- encodeString(levels(x), quote = ifelse(quote, 
            "\"", "")))
        colsep <- if (ord) 
            " < "
        else " "
        T0 <- "Levels: "
        if (is.logical(maxl)) 
            maxl <- {
                width <- width - (nchar(T0, "w") + 3L + 1L + 
                  3L)
                lenl <- cumsum(nchar(lev, "w") + nchar(colsep, 
                  "w"))
                if (n <= 1L || lenl[n] <= width) 
                  n
                else max(1L, which.max(lenl > width) - 1L)
            }
        drop <- n > maxl
        cat(if (drop) 
            paste(format(n), ""), T0, paste(if (drop) 
            c(lev[1L:max(1, maxl - 1)], "...", if (maxl > 1) lev[n])
        else lev, collapse = colsep), "\n", sep = "")
    }
    if (!isTRUE(val <- .valid.factor(x))) 
        warning(val)
    invisible(x)
}
<bytecode: 0x75f7c06c80>
<environment: namespace:base>

Generics are everywhere

Many of the base R functions you use regularly are S3 generics whose only job is to dispatch on class,

mean
function (x, ...) 
UseMethod("mean")
<bytecode: 0x75fe471738>
<environment: namespace:base>
summary
function (object, ...) 
UseMethod("summary")
<bytecode: 0x75f80fca88>
<environment: namespace:base>
plot
function (x, y, ...) 
UseMethod("plot")
<bytecode: 0x75f694dc08>
<environment: namespace:base>
t.test
function (x, ...) 
UseMethod("t.test")
<bytecode: 0x75ff1b1150>
<environment: namespace:stats>

Other generics, such as sum, dispatch without an explicit call to UseMethod():

sum
function (..., na.rm = FALSE)  .Primitive("sum")

What is S3?


S3 is R’s first and simplest OO system. S3 is R’s first and simplest OO system. S3 is informal and ad hoc, but there is a certain elegance in its minimalism: you can’t take away any part of it and still have a useful OO system.

  • Hadley Wickham, Advanced R

Dispatch

S3 dispatch is very simple - a generic calls UseMethod("func"), which searches for a function named <func>.<class>, trying each class of the first argument in order and falling back to <func>.default if none is found.

methods() lists the methods available for a generic, or all the methods available for a class,

methods("summary")
 [1] summary,ANY-method                   summary,diagonalMatrix-method       
 [3] summary,sparseMatrix-method          summary.aov                         
 [5] summary.aovlist*                     summary.aspell*                     
 [7] summary.check_packages_in_dir*       summary.connection                  
 [9] summary.data.frame                   summary.Date                        
[11] summary.default                      summary.difftime                    
[13] summary.ecdf*                        summary.factor                      
[15] summary.free1way*                    summary.glm                         
[17] summary.infl*                        summary.lm                          
[19] summary.loess*                       summary.manova                      
[21] summary.matrix                       summary.mlm*                        
[23] summary.nls*                         summary.packageStatus*              
[25] summary.pandas.core.frame.DataFrame* summary.pandas.core.series.Series*  
[27] summary.pandas.DataFrame*            summary.pandas.Series*              
[29] summary.POSIXct                      summary.POSIXlt                     
[31] summary.ppr*                         summary.prcomp*                     
[33] summary.princomp*                    summary.proc_time                   
[35] summary.python.builtin.object*       summary.rlang_error*                
[37] summary.rlang_message*               summary.rlang_trace*                
[39] summary.rlang_warning*               summary.rlang:::list_of_conditions* 
[41] summary.shingle*                     summary.srcfile                     
[43] summary.srcref                       summary.stepfun                     
[45] summary.stl*                         summary.table                       
[47] summary.trellis*                     summary.tukeysmooth*                
[49] summary.vctrs_sclr*                  summary.vctrs_vctr*                 
[51] summary.warnings                    
see '?methods' for accessing help and source code
methods(class = "factor")
 [1] [             [[            [[<-          [<-           all.equal    
 [6] Arith         as.character  as.data.frame as.Date       as.list      
[11] as.logical    as.POSIXlt    as.vector     c             cbind2       
[16] coerce        Compare       droplevels    format        free1way     
[21] initialize    is.na<-       kronecker     length<-      levels<-     
[26] Logic         Math          Ops           plot          print        
[31] rbind2        relevel       relist        rep           show         
[36] slotsFromS3   summary       Summary       xtfrm        
see '?methods' for accessing help and source code

Adding methods

Because dispatch is by name, adding a method for a new class just requires defining a new function with the proper name - no registration is needed.

( x = structure(
    c(1, 2, 3),
    class = "class_A") )
[1] 1 2 3
attr(,"class")
[1] "class_A"
( y = structure(
    c(6, 5, 4),
    class = "class_B") )
[1] 6 5 4
attr(,"class")
[1] "class_B"
print.class_A = function(x, ...) {
  cat("(Class A) ")
  print.default(unclass(x))
}
print.class_B = function(x, ...) {
  cat("(Class B) ")
  print.default(unclass(x))
}
print(x)
(Class A) [1] 1 2 3
print(y)
(Class B) [1] 6 5 4
class(x) = "class_B"
print(x)
(Class B) [1] 1 2 3
class(y) = "class_A"
print(y)
(Class A) [1] 6 5 4

Writing a good print method

In general, methods should match the generic’s signature (x, ...). For print you should then manage output via cat() or print(), and return the input invisibly so that print(x) can be used in a pipeline.

pct = structure(c(0.12, 0.5, 1), class = "percent")

print.percent = function(x, ...) {
  print(paste0(unclass(x) * 100, "%"), quote = FALSE)
  invisible(x)
}
pct
[1] 12%  50%  100%
z = print(pct)
[1] 12%  50%  100%
unclass(z)
[1] 0.12 0.50 1.00

Defining a new generic

A generic is just a function that calls UseMethod(). By convention the first argument is the object dispatched on and ... is included so that methods can add arguments.

shuffle = function(x, ...) {
  UseMethod("shuffle")
}
shuffle.default = function(x, ...) {
  stop("Class ", class(x), " is not supported by shuffle.", call. = FALSE)
}
shuffle.factor = function(x, ...) {
  factor( sample(as.character(x)), levels = sample(levels(x)) )
}
shuffle.integer = function(x, ...) {
  sample(x)
}

Shuffle results

shuffle( 1:10 )
 [1]  6  8  9  3  4  1  2  5  7 10
shuffle( factor(c("A", "B", "C", "A")) )
[1] C A A B
Levels: A C B
shuffle( c(1, 2, 3, 4, 5) )
Error:
! Class numeric is not supported by shuffle.
shuffle( letters[1:5] )
Error:
! Class character is not supported by shuffle.
methods("shuffle")
[1] shuffle.default shuffle.factor  shuffle.integer
see '?methods' for accessing help and source code

Implicit classes

Objects without a class attribute dispatch on an implicit class vector that is more detailed than what class() reports.

report = function(x) {
  UseMethod("report")
}
report.default = function(x) paste0("Class ", class(x), " does not have a method defined.")
report.integer = function(x) "I'm an integer!"
report.double  = function(x) "I'm a double!"
report.numeric = function(x) "I'm a numeric!"
report(1)
[1] "I'm a double!"
report(1L)
[1] "I'm an integer!"
report("1")
[1] "Class character does not have a method defined."
rm(report.integer, report.double)
report(1)
[1] "I'm a numeric!"
report(1L)
[1] "I'm a numeric!"
rm(report.numeric)
report(1)
[1] "Class numeric does not have a method defined."
report(1L)
[1] "Class integer does not have a method defined."

Why?

From UseMethod’s R documentation:

If the object does not have a class attribute, it has an implicit class. Matrices and arrays have class “matrix” or “array” followed by the class of the underlying vector. Most vectors have class the result of mode(x), except that integer vectors have class c("integer", "numeric") and real vectors have class c("double", "numeric").

The implicit class vector can be inspected with .class2(),

.class2(1)
[1] "double"  "numeric"
.class2(1L)
[1] "integer" "numeric"
.class2(matrix(1:4, 2))
[1] "matrix"  "array"   "integer" "numeric"