Functional programming
& rectangling

Lecture 10

Dr. Colin Rundel

Functional programming

Functions as objects

Functions are first-class objects in both languages - they can be assigned to names, stored in lists and dicts, and passed to other functions.

f = function(x) x^2
g = f
g(3)
[1] 9
l = list(sq = f, root = sqrt)
l$sq(3)
[1] 9
l[[2]](9)
[1] 3
l[2](3)
Error:
! attempt to apply non-function
def f(x):
    return x**2

g = f
g(3)
9
import math
d = {"sq": f, "root": math.sqrt}
d["sq"](3)
9
d["root"](9)
3.0

Functions as arguments

A function that takes a function as an argument is a higher-order function. Note that the function object is passed (sum), not a call to it (sum()).

do_calc = function(v, func) {
  func(v)
}
do_calc(1:3, sum)
[1] 6
do_calc(1:3, mean)
[1] 2
def do_calc(v, func):
    return func(v)
do_calc([1, 2, 3], sum)
6
do_calc([1, 2, 3], max)
3

Many built-in functions work this way:

integrate(dnorm, -1.96, 1.96)
0.9500042 with absolute error < 1e-11
sorted(["bb", "a", "ccc"], key=len)
['a', 'bb', 'ccc']

Anonymous functions

Short single-use functions without a name: function(x) or \(x) in R (4.1+), lambda in Python - a single expression, whose value is returned.

(\(x) x^2)(4)
[1] 16
integrate(\(x) sin(x)^2, 0, pi)
1.570796 with absolute error < 1.7e-14
(lambda x: x**2)(4)
16
sorted(
  ["Bob", "alice", "Carol"],
  key=lambda s: s.lower()
)
['alice', 'Bob', 'Carol']

Partial application

Partial application fixes some of a function’s arguments, producing a new function with a reduced signature. A wrapper does this by hand; partial() does it directly, by name or by position.

pow = function(base, exp) base^exp
cube = partial(pow, exp = 3)
cube(2)
[1] 8
partial(pow, 2)(5)
[1] 32
partial(pow, ... = , 3)(2)
[1] 8
from functools import partial,  Placeholder
cube = partial(pow, exp=3)
cube(2)
8
partial(pow, 2)(5)
32
partial(pow, Placeholder, 3)(2)
8

Loops vs maps

Of course, someone has to write loops. It doesn’t have to be you.

— Jenny Bryan

Most loops compute something per element and collect the results. A map does the looping and collecting, leaving only the computation. Generally not faster, just clearer than an explicit loop.

x = c(1, 4, 9)
res = numeric(length(x))
for (i in seq_along(x)) 
  res[i] = sqrt(x[i])
res
[1] 1 2 3
lapply(x, sqrt) |> str()
List of 3
 $ : num 1
 $ : num 2
 $ : num 3
sapply(x, sqrt) |> str()
 num [1:3] 1 2 3
x = [1, 4, 9]
res = []
for v in x:
  res.append(v**0.5)
res
[1.0, 2.0, 3.0]
[v**0.5 for v in x]
[1.0, 2.0, 3.0]
list(map(lambda v: v**0.5, x))
[1.0, 2.0, 3.0]

lapply() and sapply()

lapply(X, FUN, ...) applies FUN to each element of X and always returns a list; extra arguments pass through ....

sapply() simplifies the result to a vector or matrix when it can (returned type depends on the data).

lapply(1:3, \(x) x^2) |> str()
List of 3
 $ : num 1
 $ : num 4
 $ : num 9
lapply(1:3, \(x, p) x^p, p = 3) |> str()
List of 3
 $ : num 1
 $ : num 8
 $ : num 27
sapply(1:3, \(x) x^2)
[1] 1 4 9
sapply(1:3, \(x) c(x, x^2))
     [,1] [,2] [,3]
[1,]    1    2    3
[2,]    1    4    9
sapply(1:3, seq) |> str()
List of 3
 $ : int 1
 $ : int [1:2] 1 2
 $ : int [1:3] 1 2 3

map()

map(.x, .f, ...) is lapply() with a consistent interface. Includes typed variants: map_lgl(), map_int(), map_dbl(), and map_chr() that return an atomic vector of or fail.

map(1:3, \(x) x^2) |> str()
List of 3
 $ : num 1
 $ : num 4
 $ : num 9
map_dbl(1:3, \(x) x^2)
[1] 1 4 9
map_chr(1:3, \(x) paste0("n", x))
[1] "n1" "n2" "n3"
map_chr(1:3, \(x) x^2)
Error in `map_chr()`:
ℹ In index: 1.
Caused by error:
! Can't coerce from a number to a string.
map_int(1:3, \(x) x / 2)
Error in `map_int()`:
ℹ In index: 1.
Caused by error:
! Can't coerce from a number to an integer.

map2(), pmap(), and imap()

map2() iterates over two inputs in parallel, pmap() over any number (including the columns of a data frame), and imap() over elements and their names or positions. All have the same typed variants as map().

map2_chr(letters[1:3], 1:3, paste0)
[1] "a1" "b2" "c3"
map2_dbl(1:3, 4:6, \(x, y) x * y)
[1]  4 10 18
pmap_dbl(
  list(1:3, 4:6, 7:9),
  \(x, y, z) x + y + z
)
[1] 12 15 18
d = tibble(x = 1:2, y = c("a", "b"))
pmap_chr(d, \(x, y) paste(x, y))
[1] "1 a" "2 b"
imap_chr(
  c(a = 1, b = 2),
  \(x, i) paste(i, x)
)
    a     b 
"a 1" "b 2" 

Extracting elements - lookups

Often .f just extracts a piece of each loop element, purrr lets a name or position be used directly.

x = list(list(name = "a", tags = list("x", "y")),
         list(name = "b"),
         list(name = "c", tags = list("z")))
map_chr(x, "name")
[1] "a" "b" "c"
map(x, "tags") |> lengths()
[1] 2 0 1
map_chr(x, list("tags", 1))
Error in `map_chr()`:
ℹ In index: 2.
Caused by error:
! Result must be length 1, not 0.
map_chr(
  x, list("tags", 1), .default = NA
)
[1] "x" NA  "z"

keep() and discard()

keep() and discard() filter a vector or list by a predicate function, returning the elements where it is TRUE or FALSE respectively. The predicate takes the same forms as .f in map().

keep(1:10, \(x) x %% 3 == 0)
[1] 3 6 9
discard(1:10, \(x) x %% 3 == 0)
[1]  1  2  4  5  7  8 10
x = list(a = 1:3, 
         b = letters[1:2],
         c = 4:6)
keep(x, is.numeric) |> str()
List of 2
 $ a: int [1:3] 1 2 3
 $ c: int [1:3] 4 5 6
discard(x, is.numeric) |> str()
List of 1
 $ b: chr [1:2] "a" "b"

reduce() and accumulate()

reduce() collapses a list to one value by applying a two-argument function cumulatively, f(f(f(x1, x2), x3), x4).

accumulate() does the same but keeps the intermediate results.

reduce(1:5, `+`)
[1] 15
accumulate(1:5, `+`)
[1]  1  3  6 10 15
x = list(1:3, 2:4, 3:5)
reduce(x, intersect)
[1] 3
dfs = list(
  tibble(id = 1:2, a = 1:2),
  tibble(id = 1:2, b = 3:4),
  tibble(id = 1:2, c = 5:6)
)
reduce(dfs, left_join, by = "id")
# A tibble: 2 × 4
     id     a     b     c
  <int> <int> <int> <int>
1     1     1     3     5
2     2     2     4     6

Functional tools
in Python

map() and filter()

Python’s built-in map(f, iterable) and filter(pred, iterable) are the analogs of purrr’s map() and keep(). Both return lazy iterators - nothing is computed until the values are consumed (list(), a for loop, sum(), …) and they can be consumed only once.

m = map(lambda x: x**2, range(5))
m
<map object at 0x11b78b040>
list(m)
[0, 1, 4, 9, 16]
list(m)
[]
list(filter(
  lambda x: x % 2 == 0, range(10)
))
[0, 2, 4, 6, 8]
list(map(pow, [1, 2, 3], [4, 5, 6]))
[1, 32, 729]

Comprehensions

Python style prefers a comprehension to map() and filter() - transformation and filter in one expression, no lambda. The same syntax builds dicts and sets, and in parentheses a generator (next slide), which sum() and any() consume below.

[x**2 for x in range(5)]
[0, 1, 4, 9, 16]
[x**2 for x in range(10) if x % 2 == 0]
[0, 4, 16, 36, 64]
[a + b for a, b in zip("ab", "cd")]
['ac', 'bd']
{x: x**2 for x in range(4)}
{0: 0, 1: 1, 2: 4, 3: 9}
{x % 3 for x in range(10)}
{0, 1, 2}
sum(x**2 for x in range(5))
30
any(x > 3 for x in range(5))
True

Generators

A generator function uses yield instead of return. Calling it runs nothing; it returns a generator that produces one value each time it is asked, pausing after each yield. A generator expression, (f(x) for x in xs), is the one-line form.

def countdown(n):
    while n > 0:
        yield n
        n -= 1
g = countdown(3)
type(g)
<class 'generator'>
next(g)
3
next(g)
2
next(g)
1
next(g)
StopIteration
list(countdown(3))
[3, 2, 1]
[x**2 for x in countdown(3)]
[9, 4, 1]

Laziness

Creating a generator does no work; each next() runs it only as far as the following yield. So a generator can be infinite, stop early, or read a file one line at a time, and chained generators form a pipeline.

def loud(xs):
    for x in xs:
        print("yielding", x)
        yield x

g = loud([1, 2, 3])
next(g)
yielding 1
1
list(g)
yielding 2
yielding 3
[2, 3]
def naturals():
    n = 1
    while True:
        yield n
        n += 1
from itertools import islice
squares = (n**2 for n in naturals())
list(islice(squares, 5))
[1, 4, 9, 16, 25]
next(s for s in squares if s > 1000)
1024

reduce() and friends

Reductions live in the standard library - functools.reduce() for a general fold, itertools.accumulate() for the running version, and sum(), min(), max(), any(), all() for common cases. operator has function versions of the operators.

from functools import reduce
from itertools import accumulate
import operator
reduce(operator.add, range(1, 6))
15
list(accumulate(range(1, 6)))
[1, 3, 6, 10, 15]
reduce(
  lambda a, b: a + b,
  [[1, 2], [3], [4, 5]]
)
[1, 2, 3, 4, 5]
list(accumulate([1, 3, 2, 5, 4], max))
[1, 3, 3, 5, 5]

Rectangling

GitHub repositories

The GitHub API’s /orgs/{org}/repos endpoint lists an organization’s repositories as an array of objects, one per repository, each with the same 83 fields (abridged here):

[
  {
    "id": 19438,
    "name": "ggplot2",
    "owner": { "login": "tidyverse", "type": "Organization", ... },
    "stargazers_count": 7002,
    "language": "R",
    "license": { "key": "other", "name": "Other", ... },
    "topics": ["data-visualisation", "r", "visualisation"],
    ...
  },
  ...
  {
    "name": "nycflights13",
    "license": null,
    "topics": [],
    ...
  },
  ...
]

Reading JSON

jsonlite::read_json() and Python’s json.load() keep the structure as is: the array becomes a list, each object a named list (R) or dict (Python) with nested objects staying nested.

repos = jsonlite::read_json(
  "data/tidyverse_repos.json"
)
import json
path = "data/tidyverse_repos.json"
with open(path) as f:
    repos = json.load(f)
length(repos)
[1] 46
repos[[1]]$name
[1] "ggplot2"
repos[[1]]$license$name
[1] "Other"
len(repos)
46
repos[0]["name"]
'ggplot2'
repos[0]["license"]["name"]
'Other'

Extracting columns - purrr

A tidy version has one row per repository and one column per variable. The direct approach builds each column with a map_*().

tibble(
  name     = map_chr(repos, "name"),
  stars    = map_int(repos, "stargazers_count"),
  n_topics = map_int(repos, \(r) length(r$topics))
)
# A tibble: 46 × 3
  name      stars n_topics
  <chr>     <int>    <int>
1 ggplot2    7002        3
2 lubridate   805        3
3 stringr     672        3
4 dplyr      5071        3
5 readr      1041        4
# ℹ 41 more rows

Extracting columns - Python

The same idea with one comprehension per column works with pd.DataFrame() and pl.DataFrame().

pl.DataFrame({
  "name":     [r["name"] for r in repos],
  "stars":    [r["stargazers_count"] for r in repos],
  "n_topics": [len(r["topics"]) for r in repos]
})
shape: (46, 3)
name stars n_topics
str i64 i64
"ggplot2" 7002 3
"lubridate" 805 3
"stringr" 672 3
… … …
"ggbot2" 38 0
"data-dict" 107 0

Missing values

Some repositories have no license, so license is NULL in R and None in Python.

NULL cannot fill a vector slot, so map_chr() fails without .default. None is an okay list element, but indexing into it fails.

map_chr(repos, c("license", "name"))
Error in `map_chr()`:
ℹ In index: 8.
Caused by error:
! Result must be length 1, not 0.
[r["license"]["name"] for r in repos]
TypeError: 'NoneType' object is not subscriptable
license = map_chr(
  repos, c("license", "name"),
  .default = NA
)
license[7:9]
[1] "Other" NA      "Other"
license = [
  r["license"]["name"] if r["license"] else None
  for r in repos
]
license[6:9]
['Other', None, 'Other']

Exercise 1

Using repos (the tidyverse GitHub repositories) with purrr in R and comprehensions in Python,

  1. Construct a data frame with two columns: each repository’s name and its primary language.

  2. Find the five repositories with the most stars (stargazers_count).

  3. Find the names of the repositories with "r" among their topics.

“List” columns

The other approach puts the whole list into the data frame and then unnests the hierarchy. A list column holds an arbitrary object per cell; polars infers a Struct (named fields) for dicts and a List for lists, pandas stores either as object.

(d = tibble(repo = repos))
# A tibble: 46 × 1
  repo             
  <list>           
1 <named list [83]>
2 <named list [83]>
3 <named list [83]>
4 <named list [83]>
5 <named list [83]>
# ℹ 41 more rows
print(pd.DataFrame({"repo": repos}))
                       repo
0   {'id': 19438, 'node_...
1   {'id': 148017, 'node...
..                      ...
44  {'id': 1060311848, '...
45  {'id': 1184566747, '...

[46 rows x 1 columns]
print(pl.DataFrame({"repo": repos}))
shape: (46, 1)
┌─────────────────────────────────┐
│ repo                            │
│ struct[83]                      │
╞═════════════════════════════════╡
│ {19438,"MDEwOlJlcG9zaXRvcnkxOT… │
│ {148017,"MDEwOlJlcG9zaXRvcnkxN… │
│ {365649,"MDEwOlJlcG9zaXRvcnkzN… │
│ …                               │
│ {1060311848,"R_kgDOPzMTKA","gg… │
│ {1184566747,"R_kgDORpsN2w","da… │
└─────────────────────────────────┘

unnest_wider()

unnest_wider() makes each element of the named lists its own column - for list columns where every element has the same names. Types are inferred from the values, NULL becomes NA, and a nested object stays as a list column.

d |> unnest_wider(repo)
# A tibble: 46 × 83
      id node_id name  full_name private owner        html_url description fork 
   <int> <chr>   <chr> <chr>     <lgl>   <list>       <chr>    <chr>       <lgl>
1 1.94e4 MDEwOl… ggpl… tidyvers… FALSE   <named list> https:/… An impleme… FALSE
2 1.48e5 MDEwOl… lubr… tidyvers… FALSE   <named list> https:/… Make worki… FALSE
3 3.66e5 MDEwOl… stri… tidyvers… FALSE   <named list> https:/… A fresh ap… FALSE
4 6.43e6 MDEwOl… dplyr tidyvers… FALSE   <named list> https:/… dplyr: A g… FALSE
5 1.17e7 MDEwOl… readr tidyvers… FALSE   <named list> https:/… Read flat … FALSE
# ℹ 41 more rows
# ℹ 74 more variables: url <chr>, forks_url <chr>, keys_url <chr>,
#   collaborators_url <chr>, teams_url <chr>, hooks_url <chr>,
#   issue_events_url <chr>, events_url <chr>, assignees_url <chr>,
#   branches_url <chr>, tags_url <chr>, blobs_url <chr>, git_tags_url <chr>,
#   git_refs_url <chr>, trees_url <chr>, statuses_url <chr>,
#   languages_url <chr>, stargazers_url <chr>, contributors_url <chr>, …

Nested objects

license is a list column of named lists (or NULL), so a second unnest_wider() spreads it. names_sep adds a prefix, since the license’s name and url would collide with the repository’s; the NULLs become rows of NA.

(lic = d |> unnest_wider(repo) |> select(name, license) |> slice(7:9))
# A tibble: 3 × 2
  name         license         
  <chr>        <list>          
1 tidyr        <named list [5]>
2 nycflights13 <NULL>          
3 rvest        <named list [5]>
lic |>
  unnest_wider(license, names_sep = "_")
# A tibble: 3 × 6
  name      license_key license_name license_spdx_id license_url license_node_id
  <chr>     <chr>       <chr>        <chr>           <lgl>       <chr>          
1 tidyr     other       Other        NOASSERTION     NA          MDc6TGljZW5zZT…
2 nycfligh… <NA>        <NA>         <NA>            NA          <NA>           
3 rvest     other       Other        NOASSERTION     NA          MDc6TGljZW5zZT…

pandas - json_normalize()

pandas has no unnest_wider(), but the same workflow is three steps: pd.json_normalize() spreads the dicts in a column into a data frame (max_level=0 leaves nested dicts as they are), join() attaches those columns, and drop() removes the original.

df = pd.DataFrame({"repo": repos})
wide = (
  df
  .join(pd.json_normalize(df["repo"], max_level=0))
  .drop(columns="repo")
)
wide[["name", "language", "license"]]
         name language                  license
0     ggplot2        R  {'key': 'other', 'na...
1   lubridate        R  {'key': 'other', 'na...
..        ...      ...                      ...
44     ggbot2        R  {'key': 'other', 'na...
45  data-dict     Rust                     None

[46 rows x 3 columns]

pandas - nested objects

license is a column of dicts (or None), so the same three steps spread it: json_normalize() the column, add_prefix() in place of names_sep so the license’s name and url do not collide with the repository’s, then join() and drop(). The Nones become rows of NaN.

lic = wide[["name", "license"]]
lic.iloc[6:9]
           name                  license
6         tidyr  {'key': 'other', 'na...
7  nycflights13                     None
8         rvest  {'key': 'other', 'na...
lic_wide = (lic
  .join(pd.json_normalize(lic["license"]).add_prefix("license_"))
  .drop(columns="license")
)
print(lic_wide.iloc[6:9].to_string())
           name license_key license_name license_spdx_id license_url   license_node_id
6         tidyr       other        Other     NOASSERTION         NaN  MDc6TGljZW5zZTA=
7  nycflights13         NaN          NaN             NaN         NaN               NaN
8         rvest       other        Other     NOASSERTION         NaN  MDc6TGljZW5zZTA=

polars - unnest()

polars’ unnest() is unnest_wider() for a Struct column: one call spreads the fields of repo into the 83 columns, and the nested license is itself a struct (or null).

repos_pl = pl.DataFrame({"repo": repos}).unnest("repo")
repos_pl.select("name", "language", "license")
shape: (46, 3)
name language license
str str struct[5]
"ggplot2" "R" {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="}
"lubridate" "R" {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="}
"stringr" "R" {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="}
… … …
"ggbot2" "R" {"other","Other","NOASSERTION",null,"MDc6TGljZW5zZTA="}
"data-dict" "Rust" null

polars - nested objects

unnest() errors on name collisions (name, url, …), so prefix the struct’s fields first. struct.field() extracts a single field instead, renamed with alias() here since the repository already has a name column. pl.json_normalize() flattens nested dicts like the pandas function.

repos_pl.with_columns(
  pl.col("license")
    .name.prefix_fields("license_")
).unnest(
  "license"
).select(
  "name", "license_name"
)
shape: (46, 2)
name license_name
str str
"ggplot2" "Other"
"lubridate" "Other"
"stringr" "Other"
… …
"ggbot2" "Other"
"data-dict" null
repos_pl.select(
  "name",
  pl.col("license")
    .struct.field("name")
    .alias("license")
)
shape: (46, 2)
name license
str str
"ggplot2" "Other"
"lubridate" "Other"
"stringr" "Other"
… …
"ggbot2" "Other"
"data-dict" null

unnest_longer() and explode()

When an element is an array of values, the tidy shape is one row per value: unnest_longer() in tidyr, explode() in pandas and polars. Each repository has zero or more topics,

d |>
  unnest_wider(repo) |>
  select(name, topics) |>
  unnest_longer(topics)
# A tibble: 96 × 2
  name      topics            
  <chr>     <chr>             
1 ggplot2   data-visualisation
2 ggplot2   r                 
3 ggplot2   visualisation     
4 lubridate date              
5 lubridate date-time         
# ℹ 91 more rows
(wide[["name", "topics"]]
  .explode("topics")
)
         name              topics
0     ggplot2  data-visualisation
0     ggplot2                   r
..        ...                 ...
44     ggbot2                 NaN
45  data-dict                 NaN

[112 rows x 2 columns]
(repos_pl
  .select("name", "topics")
  .explode("topics")
)
shape: (112, 2)
name topics
str str
"ggplot2" "data-visualisation"
"ggplot2" "r"
"ggplot2" "visualisation"
… …
"ggbot2" null
"data-dict" null

hoist()

hoist() pulls a few fields out of a list column by name, position, or path (NA when missing) and leaves the rest. Effective when we don’t want all the columns from unnesting.

tibble(repo = repos) |>
  hoist(
    repo, name = "name",
    license = c("license", "name"),
    first_topic = list("topics", 1)
  )
# A tibble: 46 × 4
  name      license first_topic        repo             
  <chr>     <chr>   <chr>              <list>           
1 ggplot2   Other   data-visualisation <named list [82]>
2 lubridate Other   date               <named list [82]>
3 stringr   Other   r                  <named list [82]>
4 dplyr     Other   data-manipulation  <named list [82]>
5 readr     Other   csv                <named list [82]>
# ℹ 41 more rows
def at(x, i):
    return x[i] if x else None
pd.DataFrame([{
  "name": r["name"],
  "license": at(r["license"], "name"),
  "first_topic": at(r["topics"], 0)
} for r in repos])
         name license         first_topic
0     ggplot2   Other  data-visualisation
1   lubridate   Other                date
..        ...     ...                 ...
44     ggbot2   Other                 NaN
45  data-dict     NaN                 NaN

[46 rows x 3 columns]

General advice

  • Same-named entries in every element: go wider with unnest_wider(), pd.json_normalize(), or polars’ unnest().

  • Arrays of varying length: go longer with unnest_longer() or explode().

  • Avoid tidyr’s plain unnest(), which expects a list column of data frames, and unnest_auto(), which guesses between wider and longer from the data.

  • Do not rectangle what you will not use - hoist(), map_*(), and comprehensions target just the fields you need.

  • Check the column types afterwards - a list (or object) column usually means inconsistent values: a scalar in some records, an array or missing in others.

Exercise 2

Most of the 83 fields are API URLs that you will never use. Using tidyr / purrr in R and pandas or polars in Python, tidy repos into a data frame with one row per repository and just the useful fields: name, description, language, stargazers_count, forks_count, the license name (missing where there is no license), and the number of topics. Check the column types when you are done.

Summary

Functional tools

Concept R Python
anonymous function \(x) x^2, function(x) x^2 lambda x: x**2
apply to each map(x, f), lapply(x, f) [f(v) for v in x], map(f, x)
typed result map_dbl(), map_chr(), … none, check it yourself
two or more inputs map2(), pmap() zip(), map(f, a, b)
index and value imap() enumerate()
filter keep(), discard() [v for v in x if p(v)], filter()
reduce reduce(), accumulate() functools.reduce(), itertools.accumulate()
extract by name map(x, "name") [d["name"] for d in x], operator.itemgetter()
missing element .default = x["k"] if x else None, d.get("k") for a missing key
side effects walk() for loop
partial application partial(f, p = 3), ... in map() functools.partial(f, p=3)
lazy sequence none, vectors are eager generator (yield), map(), zip()

Rectangling

Task tidyr / purrr pandas polars
read JSON jsonlite::read_json() json.load() json.load()
list column tibble(x = lst) pd.DataFrame({"x": lst}), object pl.DataFrame({"x": lst}), Struct
object to columns unnest_wider() pd.json_normalize() unnest(), pl.json_normalize()
array to rows unnest_longer() explode() explode()
a few fields hoist(), map_chr(x, c("a", "b")) [d["a"]["b"] for d in x] struct.field()
missing values NULL, NA after unnesting None, NaN null

Takeaways

  • Functions are values in both languages - pass them to other functions, store them, fix some of their arguments with partial(), and write short ones inline with \(x) or lambda.

  • purrr’s map_*() family replaces lapply() / sapply() with type-stable iteration, plus shorthands for extraction, multiple inputs, filtering, and reducing. Python has map(), filter(), functools.reduce(), and lazy generators, but a comprehension is the idiom.

  • JSON becomes nested lists (R) or dicts and lists (Python). Rectangling is extracting fields and unnesting - wider for objects, longer for arrays - until each row is an observation and each column a variable.

  • Prefer targeted extraction (map_*(), hoist(), a comprehension) over unnesting everything, and check the column types when you are done.

Example


Rectangling data/discog.json