[[1]]
[1] "A"
[[2]]
[1] 0.5 1.0 1.5 2.0
[[3]]
[[3]][[1]]
[1] 1
[[4]]
function (..., na.rm = FALSE) .Primitive("sum")
Lecture 05
Lists are R’s other vector type (generic). Unlike atomic vectors they are heterogeneous - each element can be any R object: atomic vectors, other lists, functions, etc.
The printed form of a list is verbose. str() gives a compact summary of any R object’s structure and is particularly useful for lists.
Lists can contain other lists, so they do not have to be flat.
This makes them a natural way of representing tree-like data (e.g. JSON).
Elements of a list (or atomic vector) can be named. Names help avoid magic numbers when accessing elements (more readable code).
A valid name starts with a letter or . (not followed by a digit) and contains only [A-Za-z0-9._].
[ vs [[R has two addition subsetting operators.
[[ extracts a single element
$ is shorthand for named lookup with [[
[[ details[[ selects by position or by name and only ever returns one element.
Vectors of length > 1 are interpreted as recursive indexing - y[[c(3, 2)]] is y[[3]][[2]].
$ subsetting$ is a shorthand for [[ with a literal name, so y$c is equivalent to y[["c"]]. It only works with lists and objects built using lists (e.g.data frames) and it partially matches names.
Assignment with [[ or $ replaces an element or adds a new one.
Assigning NULL removes an element.
Combining an atomic vector with a list via c() produces a list (the more generic type).
unlist() goes the other way, flattening a list into an atomic vector with the usual coercion rules.
Represent the following JSON data as a list in R (we will revisit this in Python next week).
Once you have the list:
Extract the fax number using $ and [[.
Add a "mobile" phone number.
Remove the age element.
Attributes are metadata attached to an R object. Some attributes are special (e.g. names, dim, dimnames, class, levels, etc.) because they modoify how the object behaves / is treated.
Most important attributes have helper functions for getting and setting (names(), dim(), class(), levels()),
Factors are how R represents categorical data - a variable with a discrete set of possible values (called levels).
A factor is just an integer vector with two attributes: levels and class.
structure()Setting attributes one at a time is clunky - structure() attaches any number of attributes to an object in a single call and is the standard way to construct objects like this.
Knowing that factors are stored as integers explains some of their more surprising behaviors,
classThe class attribute adds a layer on top of R’s type hierarchy - previously we saw typeof() and mode().
| value | typeof() |
mode() |
class() |
|---|---|---|---|
TRUE |
logical | logical | logical |
1 |
double | numeric | numeric |
1L |
integer | numeric | integer |
"A" |
character | character | character |
NULL |
NULL | NULL | NULL |
list(1, "A") |
list | list | list |
factor("A") |
integer | numeric | factor |
matrix(1:4, 2) |
integer | numeric | matrix, array |
function(x) x^2 |
closure | function | function |
sum |
builtin | function | function |
print?print does no printing itself - UseMethod() looks at the class of x and calls the matching method. For a factor that is print.factor(), for everything without a more specific method print.default() is used.
function (x, digits = NULL, quote = TRUE, na.print = NULL, print.gap = NULL,
right = FALSE, max = NULL, width = NULL, useSource = TRUE,
...)
{
args <- pairlist(digits = digits, quote = quote, na.print = na.print,
print.gap = print.gap, right = right, max = max, width = width,
useSource = useSource, ...)
missings <- c(missing(digits), missing(quote), missing(na.print),
missing(print.gap), missing(right), missing(max), missing(width),
missing(useSource))
.Internal(print.default(x, args, missings))
}
<bytecode: 0x75fcd11968>
<environment: namespace:base>
function (x, quote = FALSE, max.levels = NULL, width = getOption("width"),
...)
{
ord <- is.ordered(x)
if (length(x) == 0L)
cat(if (ord)
"ordered"
else "factor", "()\n", sep = "")
else {
xx <- character(length(x))
xx[] <- as.character(x)
keepAttrs <- setdiff(names(attributes(x)), c("levels",
"class"))
attributes(xx)[keepAttrs] <- attributes(x)[keepAttrs]
print(xx, quote = quote, ...)
}
maxl <- max.levels %||% TRUE
if (maxl) {
n <- length(lev <- encodeString(levels(x), quote = ifelse(quote,
"\"", "")))
colsep <- if (ord)
" < "
else " "
T0 <- "Levels: "
if (is.logical(maxl))
maxl <- {
width <- width - (nchar(T0, "w") + 3L + 1L +
3L)
lenl <- cumsum(nchar(lev, "w") + nchar(colsep,
"w"))
if (n <= 1L || lenl[n] <= width)
n
else max(1L, which.max(lenl > width) - 1L)
}
drop <- n > maxl
cat(if (drop)
paste(format(n), ""), T0, paste(if (drop)
c(lev[1L:max(1, maxl - 1)], "...", if (maxl > 1) lev[n])
else lev, collapse = colsep), "\n", sep = "")
}
if (!isTRUE(val <- .valid.factor(x)))
warning(val)
invisible(x)
}
<bytecode: 0x75f7c06c80>
<environment: namespace:base>
Many of the base R functions you use regularly are S3 generics whose only job is to dispatch on class,
S3 is R’s first and simplest OO system. S3 is R’s first and simplest OO system. S3 is informal and ad hoc, but there is a certain elegance in its minimalism: you can’t take away any part of it and still have a useful OO system.
- Hadley Wickham, Advanced R
S3 dispatch is very simple - a generic calls UseMethod("func"), which searches for a function named <func>.<class>, trying each class of the first argument in order and falling back to <func>.default if none is found.
methods() lists the methods available for a generic, or all the methods available for a class,
[1] summary,ANY-method summary,diagonalMatrix-method
[3] summary,sparseMatrix-method summary.aov
[5] summary.aovlist* summary.aspell*
[7] summary.check_packages_in_dir* summary.connection
[9] summary.data.frame summary.Date
[11] summary.default summary.difftime
[13] summary.ecdf* summary.factor
[15] summary.free1way* summary.glm
[17] summary.infl* summary.lm
[19] summary.loess* summary.manova
[21] summary.matrix summary.mlm*
[23] summary.nls* summary.packageStatus*
[25] summary.pandas.core.frame.DataFrame* summary.pandas.core.series.Series*
[27] summary.pandas.DataFrame* summary.pandas.Series*
[29] summary.POSIXct summary.POSIXlt
[31] summary.ppr* summary.prcomp*
[33] summary.princomp* summary.proc_time
[35] summary.python.builtin.object* summary.rlang_error*
[37] summary.rlang_message* summary.rlang_trace*
[39] summary.rlang_warning* summary.rlang:::list_of_conditions*
[41] summary.shingle* summary.srcfile
[43] summary.srcref summary.stepfun
[45] summary.stl* summary.table
[47] summary.trellis* summary.tukeysmooth*
[49] summary.vctrs_sclr* summary.vctrs_vctr*
[51] summary.warnings
see '?methods' for accessing help and source code
[1] [ [[ [[<- [<- all.equal
[6] Arith as.character as.data.frame as.Date as.list
[11] as.logical as.POSIXlt as.vector c cbind2
[16] coerce Compare droplevels format free1way
[21] initialize is.na<- kronecker length<- levels<-
[26] Logic Math Ops plot print
[31] rbind2 relevel relist rep show
[36] slotsFromS3 summary Summary xtfrm
see '?methods' for accessing help and source code
Because dispatch is by name, adding a method for a new class just requires defining a new function with the proper name - no registration is needed.
print methodIn general, methods should match the generic’s signature (x, ...). For print you should then manage output via cat() or print(), and return the input invisibly so that print(x) can be used in a pipeline.
A generic is just a function that calls UseMethod(). By convention the first argument is the object dispatched on and ... is included so that methods can add arguments.
Objects without a class attribute dispatch on an implicit class vector that is more detailed than what class() reports.
From UseMethod’s R documentation:
If the object does not have a class attribute, it has an implicit class. Matrices and arrays have class “matrix” or “array” followed by the class of the underlying vector. Most vectors have class the result of
mode(x), except that integer vectors have classc("integer", "numeric")and real vectors have classc("double", "numeric").
Sta 523 - Fall 2026