Text Processing &
Regular Expressions

Lecture 12

Dr. Colin Rundel

Strings

Strings in R and Python

In R a string is a single element of a character vector, and string functions are vectorized over those elements.

In Python a str is an immutable sequence of characters that can be indexed and sliced, and its methods act on one string at a time.

x = c("apple", "banana", "kiwi")
nchar(x)
[1] 5 6 4
substr(x, 1, 3)
[1] "app" "ban" "kiw"
toupper(x)
[1] "APPLE"  "BANANA" "KIWI"  
x = "banana"
len(x)
6
x[0:3]
'ban'
x.upper()
'BANANA'
[
  s.upper()
  for s in ["apple", "banana", "kiwi"]
]
['APPLE', 'BANANA', 'KIWI']

stringr

stringr is the tidyverse’s string package, built on stringi and the ICU C library. Base R has equivalents for most of its functions (paste(), substr(), nchar(), gsub(), …), but their names and argument orders are inconsistent.

  • Most string-manipulation functions start with str_.

  • Most take the string as the first argument, making them convenient to use with pipes.

  • Functions are vectorized over the string inputs. Some, such as str_detect(), also accept a vector of patterns.

  • Missing strings generally produce missing results.

String operations

Most everyday string tasks have a direct equivalent in each language.

Task R / stringr Python
length str_length() len()
substring str_sub() s[i:j]
combine str_c(), str_flatten() +, sep.join()
literal split str_split(s, fixed(sep)) s.split(sep)
regex split str_split(s, pattern) re.split(pattern, s)
trim whitespace str_trim(), str_squish() s.strip(), " ".join(s.split())
pad str_pad() s.rjust(), s.ljust(), s.zfill()
case str_to_upper(), str_to_title() s.upper(), s.title()
templates str_glue() f-strings
number formatting formatC(), format() format spec, f"{x:.3f}"
many strings vectorized comprehension, .str accessor

Strings in data frames

stringr functions are vectorized, so they work directly inside mutate().

Python’s string methods are not, so pandas and polars expose vectorized versions through the .str accessor.

d = tibble(
  name = c("Bob Smith", "Alice Jones")
)
import polars as pl
d = pl.DataFrame({
  "name": ["Bob Smith", "Alice Jones"]
})
d |>
  mutate(
    upper = str_to_upper(name),
    n = str_length(name)
  )
# A tibble: 2 × 3
  name        upper           n
  <chr>       <chr>       <int>
1 Bob Smith   BOB SMITH       9
2 Alice Jones ALICE JONES    11
d.with_columns(
  upper = pl.col("name").str.to_uppercase(),
  n = pl.col("name").str.len_chars()
)
shape: (2, 3)
name upper n
str str u32
"Bob Smith" "BOB SMITH" 9
"Alice Jones" "ALICE JONES" 11

Regular expressions

Regular expressions

A regular expression (regex) is a sequence of characters that describes a pattern to be matched in text. Regexes are used to detect, extract, replace, and split strings, and to validate input.

  • The idea dates to the 1950s, when Stephen Kleene formalized regular languages.

  • They came into common use in the late 1960s and early 1970s through Unix text processing tools like ed, grep, and sed.

  • Most modern languages use a syntax descended from Perl’s. The core is shared, but each engine differs in its details.

  • In R, stringr uses the ICU engine and base R uses TRE by default or PCRE when perl = TRUE.

  • We will work through the syntax using stringr, and then see how it translates to Python’s re module.

stringr regular expression functions

Function Description
str_detect Detect the presence or absence of a pattern in a string.
str_subset Subset a vector of strings based on the presence of a pattern.
str_count Count the number of matches in a string.
str_locate Locate the first position of a pattern and return a matrix with start and end.
str_extract Extract the text corresponding to the first match.
str_match Extract capture groups formed by () from the first match.
str_split Split a string into pieces and return a list of character vectors.
str_replace Replace the first matched pattern and return a character vector.
str_remove Remove the first matched pattern and return a character vector.
str_view Show the matches made by a pattern.

Many of these functions have variants with an _all suffix (e.g. str_replace_all) which will match more than one occurrence of the pattern in a string.

Simple pattern detection

text = c("The quick brown", "fox jumps over", "the lazy dog")
str_detect(text, "quick")
[1]  TRUE FALSE FALSE
str_subset(text, "o")
[1] "The quick brown" "fox jumps over"  "the lazy dog"   
str_detect(text, "the")
[1] FALSE FALSE  TRUE
str_detect(text, regex("the", ignore_case = TRUE))
[1]  TRUE FALSE  TRUE

Escape characters

An escape character changes the interpretation of the character(s) that follow it. In R string literals the escape character is \.

Literal Character
\' single quote
\" double quote
\\ backslash
\n new line
\r carriage return
\t tab
\b backspace
\f form feed

Escapes in practice

print() shows a string as it would be typed, escapes included. cat() shows the actual characters.

print("a\tb")
[1] "a\tb"
cat("a\tb")
a   b
nchar("a\tb")
[1] 3
print("a\\b")
[1] "a\\b"
cat("a\\b")
a\b
nchar("a\\b")
[1] 3
print("a\nb")
[1] "a\nb"
cat("a\nb")
a
b
nchar("a\nb")
[1] 3
print("a\"b")
[1] "a\"b"
cat("a\"b")
a"b
nchar("a\"b")
[1] 3

Raw strings

A raw string turns off escape processing, so a backslash is just a backslash. As of R 4.0 these are written as r"(...)".

cat("\\int_0^\\infty 1/e^x")
\int_0^\infty 1/e^x
cat(r"(\int_0^\infty 1/e^x)")
\int_0^\infty 1/e^x
r"(\d+\.\d+)"
[1] "\\d+\\.\\d+"
"\\d+\\.\\d+" == r"(\d+\.\d+)"
[1] TRUE

Regex metacharacters

The power of regular expressions comes from metacharacters, which modify how the pattern is matched.

. ^ $ * + ? { } [ ] \ | ( )

To match one of these literally it needs to be escaped with a \. Regex escapes sit on top of string escapes, so without raw strings we need two levels of escaping.

To match Regex R literal R raw
. \. "\\." r"(\.)"
? \? "\\?" r"(\?)"
+ \+ "\\+" r"(\+)"
\ \\ "\\\\" r"(\\)"

Matching metacharacters

str_detect("abc[def", "\[")
## Error: '\[' is an unrecognized escape
str_detect("abc[def", "\\[")
[1] TRUE
cat("abc\\def")
abc\def
str_detect("abc\\def", "\\\\")
[1] TRUE
str_detect("abc\\def", r"(\\)")
[1] TRUE

When the pattern is plain text, fixed() treats it literally, so regex escaping is unnecessary. Normal R string-literal escaping still applies unless we use a raw string.

str_detect("abc[def", fixed("["))
[1] TRUE

XKCD’s take

Anchors

Sometimes we want a pattern to match only at a particular location in a string. Anchors match a position rather than a character.

Regex Meaning in stringr
^ Start of string (or line, in multiline mode)
$ End of string (or line, in multiline mode)
\b Word boundary
\B Not word boundary
Regex Meaning in stringr
\A Start of string, always
\Z End of string or before a trailing newline
\z End of string, always
text = "the brown owl ate the egg"
str_replace(text, "^the", "---")
[1] "--- brown owl ate the egg"
str_replace(text, "the$", "---")
[1] "the brown owl ate the egg"
str_replace(text, "egg$", "---")
[1] "the brown owl ate the ---"
str_replace_all(text, "e", "*")
[1] "th* brown owl at* th* *gg"
str_replace_all(text, "e\\b", "*")
[1] "th* brown owl at* th* egg"
str_replace_all(text, "e\\B", "*")
[1] "the brown owl ate the *gg"

Alternation

If there is more than one pattern we would like to match we can use the or (|) metacharacter. It has the lowest precedence of any operator, so use parentheses to limit its scope.

str_replace_all(text, "the|egg", "---")
[1] "--- brown owl ate --- ---"
str_replace_all(text, "a|e|i|o|u", "*")
[1] "th* br*wn *wl *t* th* *gg"
str_replace_all(text, "\\ba|e|i|o|u", "*")
[1] "th* br*wn *wl *t* th* *gg"
str_replace_all(text, "\\b(a|e|i|o|u)", "*")
[1] "the brown *wl *te the *gg"

Classes and ranges

A character class matches any one character from a set, which is written inside square brackets. Ranges simplify classes made up of contiguous letters or digits.

Class Type
[abc] Class (a or b or c)
[^abc] Negated class (not a or b or c)
[a-c] Range (lower case letter from a to c)
[A-C] Range (upper case letter from A to C)
[0-7] Range (digit from 0 to 7)
text = c("apple", "(219) 733-8965")
str_replace_all(text, "[aeiou]", "*")
[1] "*ppl*"          "(219) 733-8965"
str_replace_all(text, "[13579]", "*")
[1] "apple"          "(2**) ***-8*6*"
str_replace_all(text, "[1-5a-f]", "*")
[1] "*ppl*"          "(**9) 7**-896*"
str_replace_all(
  text, "[^1-5a-f]", "*"
)
[1] "a***e"          "*21****33****5"

Character classes

Some classes are common enough that they have built-in shorthands.

Meta char POSIX class Description
. Any character except a line terminator, unless dotall = TRUE
\s [[:space:]] White space
\S Not white space
\d [[:digit:]] Unicode decimal digit; use [0-9] for ASCII digits
\D Not digit
\w Unicode word character, including letters, digits, combining marks, and _
\W Not word
[[:alpha:]] Letter
[[:punct:]] Punctuation

Unicode character classes

stringr’s character classes follow Unicode categories, so they include characters beyond ASCII. Here ٣ is the Arabic-Indic digit three, _ is both a word character and punctuation, and + is a symbol rather than punctuation.

tibble(
  char = c("A", "é", "٣", "_", "+"),
  digit = str_detect(char, r"(\d)"),
  word = str_detect(char, r"(\w)"),
  punctuation = str_detect(char, "[[:punct:]]")
)
# A tibble: 5 × 4
  char  digit word  punctuation
  <chr> <lgl> <lgl> <lgl>      
1 A     FALSE TRUE  FALSE      
2 é     FALSE TRUE  FALSE      
3 ٣     TRUE  TRUE  FALSE      
4 _     FALSE TRUE  TRUE       
5 +     FALSE FALSE FALSE      

Example

How would we write a regular expression to match a telephone number with the form (###) ###-####?

text = c("apple", "(219) 733-8965", "(329) 293-8753")
str_detect(text, "(\\d\\d\\d) \\d\\d\\d-\\d\\d\\d\\d")
[1] FALSE FALSE FALSE
str_detect(text, "\\(\\d\\d\\d\\) \\d\\d\\d-\\d\\d\\d\\d")
[1] FALSE  TRUE  TRUE

Quantifiers

Attached to literals, character classes, or groups, these allow a match to repeat some number of times.

Quantifier Description
* Match 0 or more
+ Match 1 or more
? Match 0 or 1
{3} Match exactly 3
{3,} Match 3 or more
{3,5} Match 3, 4 or 5

Example

How would we improve our previous regular expression for matching a telephone number with the form (###) ###-####?

text = c("apple", "(219) 733-8965", "(329) 293-8753")
str_detect(text, "\\(\\d\\d\\d\\) \\d\\d\\d-\\d\\d\\d\\d")
[1] FALSE  TRUE  TRUE
str_detect(text, "\\(\\d{3}\\) \\d{3}-\\d{4}")
[1] FALSE  TRUE  TRUE
str_extract(text, "\\(\\d{3}\\) \\d{3}-\\d{4}")
[1] NA               "(219) 733-8965" "(329) 293-8753"

Exercise 1

For this exercise, treat the first token as the first name and everything after the first space as the last name, including prefixes and compound names. Match vowels (a, e, i, o, u) without regard to case.

For the following vector of randomly generated names, write a regular expression that,

  • detects if the person’s first name starts with a vowel (a,e,i,o,u)

  • detects if the person’s last name starts with a vowel

  • detects if either the person’s first or last name starts with a vowel

  • detects if neither the person’s first nor last name starts with a vowel

x = c("Jeremy Cruz", "Nathaniel Le", "Jasmine Chu", "Bradley Calderon Raygoza",
      "Quinten Weller", "Katelien Kanamu-Hauanio", "Zuhriyaa al-Amen",
      "Travale York", "Alexis Ahmed", "David Alcocer", "Jairo Martinez",
      "Dwone Gallegos", "Amanda Sherwood", "Hadiyya el-Eid", "Shaimaaa al-Can",
      "Sarah Love", "Shelby Villano", "Sundus al-Hashmi", "Dyani Loving",
      "Shanelle Douglas")

Extracting matches

str_extract() returns the first match in each string, or NA when there is none. str_extract_all() returns every match, as a list with one element for each string.

x = c(
  "Midterms on 2026-10-14 and 2026-11-18",
  "Final on 2026-12-12",
  "No project"
)
date = "\\d{4}-\\d{2}-\\d{2}"
str_extract(x, date)
[1] "2026-10-14" "2026-12-12" NA          
str_extract_all(x, date)
[[1]]
[1] "2026-10-14" "2026-11-18"

[[2]]
[1] "2026-12-12"

[[3]]
character(0)

Greedy vs non-greedy matching

At a given starting position, a greedy quantifier tries to consume as much text as possible while allowing the rest of the pattern to match. Adding ? makes it non-greedy, so it consumes as little as possible while still allowing the rest to match.

text = "<div class='main'> <div> <a href='here.pdf'>Here!</a> </div> </div>"
str_extract(text, "<div>.*</div>")
[1] "<div> <a href='here.pdf'>Here!</a> </div> </div>"
str_extract(text, "<div>.*?</div>")
[1] "<div> <a href='here.pdf'>Here!</a> </div>"

Groups

Groups allow you to connect pieces of a regular expression for modification or capture.

Group Description
(a|b) match literal “a” or “b”, group either
a(bc)? match “a” or “abc”, group bc or nothing
(abc)def(ghi) match “abcdefghi”, group abc and ghi
(?:abc) match “abc”, non-capturing group

Groups in stringr

str_extract() returns the full match by default. Its group argument can select a capturing group. str_match() returns a matrix with the full match in the first column, followed by one column per capturing group.

text = c("Bob Smith", "Alice Smith", "Apple")
str_extract(text, "^(\\w+) \\w+")
[1] "Bob Smith"   "Alice Smith" NA           
str_extract(text, "^(\\w+) \\w+", group = 1)
[1] "Bob"   "Alice" NA     
str_match(text, "^(\\w+) \\w+")
     [,1]          [,2]   
[1,] "Bob Smith"   "Bob"  
[2,] "Alice Smith" "Alice"
[3,] NA            NA     
str_extract(text, "^(\\w+) (\\w+)")
[1] "Bob Smith"   "Alice Smith" NA           
str_match(text, "^(\\w+) (\\w+)")
     [,1]          [,2]    [,3]   
[1,] "Bob Smith"   "Bob"   "Smith"
[2,] "Alice Smith" "Alice" "Smith"
[3,] NA            NA      NA     

Multiple matches and groups

str_match_all() returns a list with one matrix per input string. Each row is a match. The first column contains the full match, followed by one column per capturing group.

str_match_all("A12 B34", "([A-Z])([0-9]+)")
[[1]]
     [,1]  [,2] [,3]
[1,] "A12" "A"  "12"
[2,] "B34" "B"  "34"

Both matches are retained, with the letter and digits captured separately. An input with no matches gets a matrix with zero rows.

Backreferences

Backreferences refer to previously captured groups using \1, \2, etc. They can be used in a replacement string, or within the pattern itself to match repeated text.

text = c("Bob Smith", "Alice Smith", "Apple")
str_replace(text, "^(\\w+) (\\w+)", "\\2, \\1")
[1] "Smith, Bob"   "Smith, Alice" "Apple"       
str_detect(c("abab", "cdcd", "abcd"), "^(..)\\1$")
[1]  TRUE  TRUE FALSE
str_extract(c("hello", "banana", "apple"), "(.)\\1")
[1] "ll" NA   "pp"

Flags

Flags change how a pattern is interpreted. In stringr they are set by wrapping the pattern in regex().

Argument Behavior
ignore_case = TRUE match without regard to case
multiline = TRUE ^ and $ match at the start and end of each line
dotall = TRUE . also matches \n
comments = TRUE ignore whitespace and # comments in the pattern
str_count("the cat\nthe hat", "^the")
[1] 1
str_count("the cat\nthe hat", regex("^the", multiline = TRUE))
[1] 2

Regex in data frames

stringr’s functions are vectorized, so they can be used directly on the columns of a data frame.

d = tibble(x = c("Bob Smith", "Alice Jones"))
d |>
  mutate(
    first = str_extract(x, "^\\w+"),
    s = str_detect(x, "S")
  )
# A tibble: 2 × 3
  x           first s    
  <chr>       <chr> <lgl>
1 Bob Smith   Bob   TRUE 
2 Alice Jones Alice FALSE
d |>
  filter(str_detect(x, "^A"))
# A tibble: 1 × 1
  x          
  <chr>      
1 Alice Jones

Exercise 2

text = c(
  "apple",
  "219 733 8965",
  "329-293-8753",
  "Work: (579) 499-7527; Home: (543) 355 3679"
)
  • Write a regular expression that will extract all of the phone numbers contained in the vector above. Expect 0, 1, 1, and 2 matches across the four inputs.

  • Once that works, use capturing groups with str_match_all() to retain every phone number while separating the area code from the remaining number. Keep the matches associated with their original input string.

Regular expressions in Python

Python’s re module

Function Description
re.search(p, s) Find the first match anywhere in the string, returns a Match or None.
re.match(p, s) Match only at the start of the string.
re.fullmatch(p, s) Match only if the entire string matches.
re.findall(p, s) Return all matches or captured groups as a list of strings or tuples.
re.finditer(p, s) Return all matches as an iterator of Match objects.
re.sub(p, repl, s) Replace all matches with repl.
re.split(p, s) Split the string at each match.
re.compile(p) Compile a pattern into an object that can be reused.

Most patterns shown so far translate directly, but regex engines differ in some details as well as their interfaces.

For example re does not support the POSIX classes like [[:digit:]]. Use \d, \s, \w, or an explicit class such as [a-z].

Raw strings in Python

Python’s raw strings are written r"...", and the convention is to write every regex pattern as one.

print("\\int_0^\\infty 1/e^x")
\int_0^\infty 1/e^x
print(r"\int_0^\infty 1/e^x")
\int_0^\infty 1/e^x
import re
text = "the brown owl ate the egg"
re.sub(r"\bow", "---", text)
'the brown ---l ate the egg'
re.sub("\bow", "---", text)
'the brown owl ate the egg'

Detecting patterns

re.search() takes a single string, so the vectorized str_detect() and str_subset() become comprehensions.

text = ["The quick brown", "fox jumps over", "the lazy dog"]
[re.search("quick", s) for s in text]
[<re.Match object; span=(4, 9), match='quick'>, None, None]
[bool(re.search("quick", s)) for s in text]
[True, False, False]
[s for s in text if re.search("o", s)]
['The quick brown', 'fox jumps over', 'the lazy dog']
[bool(re.search("the", s)) for s in text]
[False, False, True]
[bool(re.search("the", s, re.IGNORECASE)) for s in text]
[True, False, True]

Match objects

re.search() returns a Match object, or None when there is no match. None is falsy, so if m: is the usual test. Positions are 0-based and the end is exclusive.

x = "fox jumps over"
m = re.search("o", x); m
<re.Match object; span=(1, 2), match='o'>
m.group()
'o'
m.start(), m.end()
(1, 2)
print(re.search("z", x))
None
re.findall("o", x)
['o', 'o']
[
  m.start()
  for m in re.finditer("o", x)
]
[1, 10]

Replacing and splitting

re.sub() plays the role of both str_replace() and str_replace_all(). Backreferences in the replacement work the same way as in stringr.

text = "the brown owl ate the egg"
re.sub(r"the", "---", text)
'--- brown owl ate --- egg'
re.sub(r"the", "---", text, count=1)
'--- brown owl ate the egg'
re.sub(r"\b(a|e|i|o|u)", "*", text)
'the brown *wl *te the *gg'
re.sub(r"^(\w+) (\w+)", r"\2, \1", "Bob Smith")
'Smith, Bob'
re.split(r"\s+", text)
['the', 'brown', 'owl', 'ate', 'the', 'egg']

Match groups in Python

A Match object gives access to the capture groups, with group(0) being the full match.

m = re.search(
  r"^(\w+) (\w+)", "Bob Smith"
)
m.group(0)
'Bob Smith'
m.group(1)
'Bob'
m.group(2)
'Smith'
m.groups()
('Bob', 'Smith')

Flags and compiled patterns

Flags are passed as an additional argument and combined with |. re.compile() stores a pattern and its flags in an object with the same methods as the re module.

Behavior stringr::regex() Python re
ignore case ignore_case = TRUE re.IGNORECASE, re.I
^ and $ match each line multiline = TRUE re.MULTILINE, re.M
. matches \n dotall = TRUE re.DOTALL, re.S
whitespace and comments comments = TRUE re.VERBOSE, re.X
text = "The cat in\nthe hat"
re.findall(r"^the", text)
[]
re.findall(r"^the", text, re.M)
['the']
re.findall(r"^the", text, re.M | re.I)
['The', 'the']
p = re.compile(r"^the", re.M | re.I)
p.findall(text)
['The', 'the']
p.sub("a", text)
'a cat in\na hat'

Regex in pandas and polars

Many .str methods in pandas and polars accept regular expressions. polars uses Rust’s regex engine, which does not support backreferences in patterns.

import pandas as pd

s = pd.Series(
  ["Bob Smith", "Alice Jones"]
)
s.str.extract(r"^(\w+) (\w+)")
       0      1
0    Bob  Smith
1  Alice  Jones
s.str.contains(r"S")
0     True
1    False
dtype: bool
d = pl.DataFrame({
  "x": ["Bob Smith", "Alice Jones"]
})
d.with_columns(
  first = pl.col("x").str.extract(r"^(\w+)"),
  s = pl.col("x").str.contains(r"S")
)
shape: (2, 3)
x first s
str str bool
"Bob Smith" "Bob" true
"Alice Jones" "Alice" false

Summary

Regex in R and Python

Task stringr Python
pattern string "\\d+", r"(\d+)" r"\d+"
detect str_detect() re.search()
subset str_subset() comprehension with re.search()
first match str_extract() m = re.search(p, s)
m.group() if m else None
all matches str_extract_all() re.findall(), re.finditer()
capture groups str_match(), str_match_all() m.group(1), m.groups(), re.findall()
replace first str_replace() re.sub(count=1)
replace all str_replace_all() re.sub()
split str_split() re.split()
no match: detect str_detect() → FALSE re.search() → None
no match: first str_extract() → NA check for None before .group()
no match: all str_extract_all() → character(0) per input re.findall() → []
literal pattern fixed() in, s.replace(), re.escape()
flags regex(ignore_case = TRUE) re.IGNORECASE
data frame columns vectorized .str.contains(), .str.extract()

Takeaways

  • stringr gives R a consistent, vectorized set of string functions. Python’s equivalents are methods on a single str, applied to many strings with a comprehension or the .str accessor of pandas and polars.

  • Regular expression syntax is largely shared between the languages. stringr generally takes the string first and works over vectors. Python’s re takes the pattern first and works on one string. Return values depend on the operation.

  • Regex escapes sit on top of string escapes. Use raw strings, r"(...)" in R and r"..." in Python, to avoid doubling every backslash.

  • Build patterns up incrementally and test them against examples that should and should not match. If a pattern becomes unreadable, a few simpler steps are usually the better choice.