[1] 9
Lecture 10
Functions are first-class objects in both languages - they can be assigned to names, stored in lists and dicts, and passed to other functions.
A function that takes a function as an argument is a higher-order function. Note that the function object is passed (sum), not a call to it (sum()).
Short single-use functions without a name: function(x) or \(x) in R (4.1+), lambda in Python - a single expression, whose value is returned.
Partial application fixes some of a function’s arguments, producing a new function with a reduced signature. A wrapper does this by hand; partial() does it directly, by name or by position.
Of course, someone has to write loops. It doesn’t have to be you.
— Jenny Bryan
Most loops compute something per element and collect the results. A map does the looping and collecting, leaving only the computation. Generally not faster, just clearer than an explicit loop.
lapply() and sapply()lapply(X, FUN, ...) applies FUN to each element of X and always returns a list; extra arguments pass through ....
sapply() simplifies the result to a vector or matrix when it can (returned type depends on the data).
map()map(.x, .f, ...) is lapply() with a consistent interface. Includes typed variants: map_lgl(), map_int(), map_dbl(), and map_chr() that return an atomic vector of or fail.
map2(), pmap(), and imap()map2() iterates over two inputs in parallel, pmap() over any number (including the columns of a data frame), and imap() over elements and their names or positions. All have the same typed variants as map().
[1] "a1" "b2" "c3"
[1] 4 10 18
Often .f just extracts a piece of each loop element, purrr lets a name or position be used directly.
keep() and discard()keep() and discard() filter a vector or list by a predicate function, returning the elements where it is TRUE or FALSE respectively. The predicate takes the same forms as .f in map().
reduce() and accumulate()reduce() collapses a list to one value by applying a two-argument function cumulatively, f(f(f(x1, x2), x3), x4).
accumulate() does the same but keeps the intermediate results.
map() and filter()Python’s built-in map(f, iterable) and filter(pred, iterable) are the analogs of purrr’s map() and keep(). Both return lazy iterators - nothing is computed until the values are consumed (list(), a for loop, sum(), …) and they can be consumed only once.
Python style prefers a comprehension to map() and filter() - transformation and filter in one expression, no lambda. The same syntax builds dicts and sets, and in parentheses a generator (next slide), which sum() and any() consume below.
[0, 1, 4, 9, 16]
[0, 4, 16, 36, 64]
A generator function uses yield instead of return. Calling it runs nothing; it returns a generator that produces one value each time it is asked, pausing after each yield. A generator expression, (f(x) for x in xs), is the one-line form.
Creating a generator does no work; each next() runs it only as far as the following yield. So a generator can be infinite, stop early, or read a file one line at a time, and chained generators form a pipeline.
reduce() and friendsReductions live in the standard library - functools.reduce() for a general fold, itertools.accumulate() for the running version, and sum(), min(), max(), any(), all() for common cases. operator has function versions of the operators.
The GitHub API’s /orgs/{org}/repos endpoint lists an organization’s repositories as an array of objects, one per repository, each with the same 83 fields (abridged here):
[
{
"id": 19438,
"name": "ggplot2",
"owner": { "login": "tidyverse", "type": "Organization", ... },
"stargazers_count": 7002,
"language": "R",
"license": { "key": "other", "name": "Other", ... },
"topics": ["data-visualisation", "r", "visualisation"],
...
},
...
{
"name": "nycflights13",
"license": null,
"topics": [],
...
},
...
]jsonlite::read_json() and Python’s json.load() keep the structure as is: the array becomes a list, each object a named list (R) or dict (Python) with nested objects staying nested.
A tidy version has one row per repository and one column per variable. The direct approach builds each column with a map_*().
The same idea with one comprehension per column works with pd.DataFrame() and pl.DataFrame().
Some repositories have no license, so license is NULL in R and None in Python.
NULL cannot fill a vector slot, so map_chr() fails without .default. None is an okay list element, but indexing into it fails.
Error in `map_chr()`:
ℹ In index: 8.
Caused by error:
! Result must be length 1, not 0.
Using repos (the tidyverse GitHub repositories) with purrr in R and comprehensions in Python,
Construct a data frame with two columns: each repository’s name and its primary language.
Find the five repositories with the most stars (stargazers_count).
Find the names of the repositories with "r" among their topics.
The other approach puts the whole list into the data frame and then unnests the hierarchy. A list column holds an arbitrary object per cell; polars infers a Struct (named fields) for dicts and a List for lists, pandas stores either as object.
# A tibble: 46 × 1
repo
<list>
1 <named list [83]>
2 <named list [83]>
3 <named list [83]>
4 <named list [83]>
5 <named list [83]>
# ℹ 41 more rows
shape: (46, 1)
┌─────────────────────────────────┐
│ repo │
│ struct[83] │
╞═════════════════════════════════╡
│ {19438,"MDEwOlJlcG9zaXRvcnkxOT… │
│ {148017,"MDEwOlJlcG9zaXRvcnkxN… │
│ {365649,"MDEwOlJlcG9zaXRvcnkzN… │
│ … │
│ {1060311848,"R_kgDOPzMTKA","gg… │
│ {1184566747,"R_kgDORpsN2w","da… │
└─────────────────────────────────┘
unnest_wider()unnest_wider() makes each element of the named lists its own column - for list columns where every element has the same names. Types are inferred from the values, NULL becomes NA, and a nested object stays as a list column.
# A tibble: 46 × 83
id node_id name full_name private owner html_url description fork
<int> <chr> <chr> <chr> <lgl> <list> <chr> <chr> <lgl>
1 1.94e4 MDEwOl… ggpl… tidyvers… FALSE <named list> https:/… An impleme… FALSE
2 1.48e5 MDEwOl… lubr… tidyvers… FALSE <named list> https:/… Make worki… FALSE
3 3.66e5 MDEwOl… stri… tidyvers… FALSE <named list> https:/… A fresh ap… FALSE
4 6.43e6 MDEwOl… dplyr tidyvers… FALSE <named list> https:/… dplyr: A g… FALSE
5 1.17e7 MDEwOl… readr tidyvers… FALSE <named list> https:/… Read flat … FALSE
# ℹ 41 more rows
# ℹ 74 more variables: url <chr>, forks_url <chr>, keys_url <chr>,
# collaborators_url <chr>, teams_url <chr>, hooks_url <chr>,
# issue_events_url <chr>, events_url <chr>, assignees_url <chr>,
# branches_url <chr>, tags_url <chr>, blobs_url <chr>, git_tags_url <chr>,
# git_refs_url <chr>, trees_url <chr>, statuses_url <chr>,
# languages_url <chr>, stargazers_url <chr>, contributors_url <chr>, …
license is a list column of named lists (or NULL), so a second unnest_wider() spreads it. names_sep adds a prefix, since the license’s name and url would collide with the repository’s; the NULLs become rows of NA.
# A tibble: 3 × 2
name license
<chr> <list>
1 tidyr <named list [5]>
2 nycflights13 <NULL>
3 rvest <named list [5]>
# A tibble: 3 × 6
name license_key license_name license_spdx_id license_url license_node_id
<chr> <chr> <chr> <chr> <lgl> <chr>
1 tidyr other Other NOASSERTION NA MDc6TGljZW5zZT…
2 nycfligh… <NA> <NA> <NA> NA <NA>
3 rvest other Other NOASSERTION NA MDc6TGljZW5zZT…
json_normalize()pandas has no unnest_wider(), but the same workflow is three steps: pd.json_normalize() spreads the dicts in a column into a data frame (max_level=0 leaves nested dicts as they are), join() attaches those columns, and drop() removes the original.
name language license
0 ggplot2 R {'key': 'other', 'na...
1 lubridate R {'key': 'other', 'na...
.. ... ... ...
44 ggbot2 R {'key': 'other', 'na...
45 data-dict Rust None
[46 rows x 3 columns]
license is a column of dicts (or None), so the same three steps spread it: json_normalize() the column, add_prefix() in place of names_sep so the license’s name and url do not collide with the repository’s, then join() and drop(). The Nones become rows of NaN.
name license
6 tidyr {'key': 'other', 'na...
7 nycflights13 None
8 rvest {'key': 'other', 'na...
name license_key license_name license_spdx_id license_url license_node_id
6 tidyr other Other NOASSERTION NaN MDc6TGljZW5zZTA=
7 nycflights13 NaN NaN NaN NaN NaN
8 rvest other Other NOASSERTION NaN MDc6TGljZW5zZTA=
unnest()polars’ unnest() is unnest_wider() for a Struct column: one call spreads the fields of repo into the 83 columns, and the nested license is itself a struct (or null).
| name | language | license |
|---|---|---|
| str | str | struct[5] |
| "ggplot2" | "R" | {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="} |
| "lubridate" | "R" | {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="} |
| "stringr" | "R" | {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="} |
| … | … | … |
| "ggbot2" | "R" | {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="} |
| "data-dict" | "Rust" | null |
unnest() errors on name collisions (name, url, …), so prefix the struct’s fields first. struct.field() extracts a single field instead, renamed with alias() here since the repository already has a name column. pl.json_normalize() flattens nested dicts like the pandas function.
unnest_longer() and explode()When an element is an array of values, the tidy shape is one row per value: unnest_longer() in tidyr, explode() in pandas and polars. Each repository has zero or more topics,
hoist()hoist() pulls a few fields out of a list column by name, position, or path (NA when missing) and leaves the rest. Effective when we don’t want all the columns from unnesting.
# A tibble: 46 × 4
name license first_topic repo
<chr> <chr> <chr> <list>
1 ggplot2 Other data-visualisation <named list [82]>
2 lubridate Other date <named list [82]>
3 stringr Other r <named list [82]>
4 dplyr Other data-manipulation <named list [82]>
5 readr Other csv <named list [82]>
# ℹ 41 more rows
name license first_topic
0 ggplot2 Other data-visualisation
1 lubridate Other date
.. ... ... ...
44 ggbot2 Other NaN
45 data-dict NaN NaN
[46 rows x 3 columns]
Same-named entries in every element: go wider with unnest_wider(), pd.json_normalize(), or polars’ unnest().
Arrays of varying length: go longer with unnest_longer() or explode().
Avoid tidyr’s plain unnest(), which expects a list column of data frames, and unnest_auto(), which guesses between wider and longer from the data.
Do not rectangle what you will not use - hoist(), map_*(), and comprehensions target just the fields you need.
Check the column types afterwards - a list (or object) column usually means inconsistent values: a scalar in some records, an array or missing in others.
Most of the 83 fields are API URLs that you will never use. Using tidyr / purrr in R and pandas or polars in Python, tidy repos into a data frame with one row per repository and just the useful fields: name, description, language, stargazers_count, forks_count, the license name (missing where there is no license), and the number of topics. Check the column types when you are done.
| Concept | R | Python |
|---|---|---|
| anonymous function | \(x) x^2, function(x) x^2 |
lambda x: x**2 |
| apply to each | map(x, f), lapply(x, f) |
[f(v) for v in x], map(f, x) |
| typed result | map_dbl(), map_chr(), … |
none, check it yourself |
| two or more inputs | map2(), pmap() |
zip(), map(f, a, b) |
| index and value | imap() |
enumerate() |
| filter | keep(), discard() |
[v for v in x if p(v)], filter() |
| reduce | reduce(), accumulate() |
functools.reduce(), itertools.accumulate() |
| extract by name | map(x, "name") |
[d["name"] for d in x], operator.itemgetter() |
| missing element | .default = |
x["k"] if x else None, d.get("k") for a missing key |
| side effects | walk() |
for loop |
| partial application | partial(f, p = 3), ... in map() |
functools.partial(f, p=3) |
| lazy sequence | none, vectors are eager | generator (yield), map(), zip() |
| Task | tidyr / purrr | pandas | polars |
|---|---|---|---|
| read JSON | jsonlite::read_json() |
json.load() |
json.load() |
| list column | tibble(x = lst) |
pd.DataFrame({"x": lst}), object |
pl.DataFrame({"x": lst}), Struct |
| object to columns | unnest_wider() |
pd.json_normalize() |
unnest(), pl.json_normalize() |
| array to rows | unnest_longer() |
explode() |
explode() |
| a few fields | hoist(), map_chr(x, c("a", "b")) |
[d["a"]["b"] for d in x] |
struct.field() |
| missing values | NULL, NA after unnesting |
None, NaN |
null |
Functions are values in both languages - pass them to other functions, store them, fix some of their arguments with partial(), and write short ones inline with \(x) or lambda.
purrr’s map_*() family replaces lapply() / sapply() with type-stable iteration, plus shorthands for extraction, multiple inputs, filtering, and reducing. Python has map(), filter(), functools.reduce(), and lazy generators, but a comprehension is the idiom.
JSON becomes nested lists (R) or dicts and lists (Python). Rectangling is extracting fields and unnesting - wider for objects, longer for arrays - until each row is an observation and each column a variable.
Prefer targeted extraction (map_*(), hoist(), a comprehension) over unnesting everything, and check the column types when you are done.
Rectangling data/discog.json
Sta 523 - Fall 2026