Lecture 12
In R a string is a single element of a character vector, and string functions are vectorized over those elements.
In Python a str is an immutable sequence of characters that can be indexed and sliced, and its methods act on one string at a time.
stringrstringr is the tidyverse’s string package, built on stringi and the ICU C library. Base R has equivalents for most of its functions (paste(), substr(), nchar(), gsub(), …), but their names and argument orders are inconsistent.
Most string-manipulation functions start with str_.
Most take the string as the first argument, making them convenient to use with pipes.
Functions are vectorized over the string inputs. Some, such as str_detect(), also accept a vector of patterns.
Missing strings generally produce missing results.
Most everyday string tasks have a direct equivalent in each language.
| Task | R / stringr |
Python |
|---|---|---|
| length | str_length() |
len() |
| substring | str_sub() |
s[i:j] |
| combine | str_c(), str_flatten() |
+, sep.join() |
| literal split | str_split(s, fixed(sep)) |
s.split(sep) |
| regex split | str_split(s, pattern) |
re.split(pattern, s) |
| trim whitespace | str_trim(), str_squish() |
s.strip(), " ".join(s.split()) |
| pad | str_pad() |
s.rjust(), s.ljust(), s.zfill() |
| case | str_to_upper(), str_to_title() |
s.upper(), s.title() |
| templates | str_glue() |
f-strings |
| number formatting | formatC(), format() |
format spec, f"{x:.3f}" |
| many strings | vectorized | comprehension, .str accessor |
stringr functions are vectorized, so they work directly inside mutate().
Python’s string methods are not, so pandas and polars expose vectorized versions through the .str accessor.
A regular expression (regex) is a sequence of characters that describes a pattern to be matched in text. Regexes are used to detect, extract, replace, and split strings, and to validate input.
The idea dates to the 1950s, when Stephen Kleene formalized regular languages.
They came into common use in the late 1960s and early 1970s through Unix text processing tools like ed, grep, and sed.
Most modern languages use a syntax descended from Perl’s. The core is shared, but each engine differs in its details.
In R, stringr uses the ICU engine and base R uses TRE by default or PCRE when perl = TRUE.
We will work through the syntax using stringr, and then see how it translates to Python’s re module.
stringr regular expression functions| Function | Description |
|---|---|
str_detect |
Detect the presence or absence of a pattern in a string. |
str_subset |
Subset a vector of strings based on the presence of a pattern. |
str_count |
Count the number of matches in a string. |
str_locate |
Locate the first position of a pattern and return a matrix with start and end. |
str_extract |
Extract the text corresponding to the first match. |
str_match |
Extract capture groups formed by () from the first match. |
str_split |
Split a string into pieces and return a list of character vectors. |
str_replace |
Replace the first matched pattern and return a character vector. |
str_remove |
Remove the first matched pattern and return a character vector. |
str_view |
Show the matches made by a pattern. |
Many of these functions have variants with an _all suffix (e.g. str_replace_all) which will match more than one occurrence of the pattern in a string.
An escape character changes the interpretation of the character(s) that follow it. In R string literals the escape character is \.
| Literal | Character |
|---|---|
\' |
single quote |
\" |
double quote |
\\ |
backslash |
\n |
new line |
\r |
carriage return |
\t |
tab |
\b |
backspace |
\f |
form feed |
print() shows a string as it would be typed, escapes included. cat() shows the actual characters.
A raw string turns off escape processing, so a backslash is just a backslash. As of R 4.0 these are written as r"(...)".
The power of regular expressions comes from metacharacters, which modify how the pattern is matched.
To match one of these literally it needs to be escaped with a \. Regex escapes sit on top of string escapes, so without raw strings we need two levels of escaping.
| To match | Regex | R literal | R raw |
|---|---|---|---|
. |
\. |
"\\." |
r"(\.)" |
? |
\? |
"\\?" |
r"(\?)" |
+ |
\+ |
"\\+" |
r"(\+)" |
\ |
\\ |
"\\\\" |
r"(\\)" |
Sometimes we want a pattern to match only at a particular location in a string. Anchors match a position rather than a character.
| Regex | Meaning in stringr |
|---|---|
^ |
Start of string (or line, in multiline mode) |
$ |
End of string (or line, in multiline mode) |
\b |
Word boundary |
\B |
Not word boundary |
| Regex | Meaning in stringr |
|---|---|
\A |
Start of string, always |
\Z |
End of string or before a trailing newline |
\z |
End of string, always |
If there is more than one pattern we would like to match we can use the or (|) metacharacter. It has the lowest precedence of any operator, so use parentheses to limit its scope.
A character class matches any one character from a set, which is written inside square brackets. Ranges simplify classes made up of contiguous letters or digits.
| Class | Type |
|---|---|
[abc] |
Class (a or b or c) |
[^abc] |
Negated class (not a or b or c) |
[a-c] |
Range (lower case letter from a to c) |
[A-C] |
Range (upper case letter from A to C) |
[0-7] |
Range (digit from 0 to 7) |
Some classes are common enough that they have built-in shorthands.
| Meta char | POSIX class | Description |
|---|---|---|
. |
Any character except a line terminator, unless dotall = TRUE |
|
\s |
[[:space:]] |
White space |
\S |
Not white space | |
\d |
[[:digit:]] |
Unicode decimal digit; use [0-9] for ASCII digits |
\D |
Not digit | |
\w |
Unicode word character, including letters, digits, combining marks, and _ |
|
\W |
Not word | |
[[:alpha:]] |
Letter | |
[[:punct:]] |
Punctuation |
stringr’s character classes follow Unicode categories, so they include characters beyond ASCII. Here ٣ is the Arabic-Indic digit three, _ is both a word character and punctuation, and + is a symbol rather than punctuation.
# A tibble: 5 × 4
char digit word punctuation
<chr> <lgl> <lgl> <lgl>
1 A FALSE TRUE FALSE
2 é FALSE TRUE FALSE
3 ٣ TRUE TRUE FALSE
4 _ FALSE TRUE TRUE
5 + FALSE FALSE FALSE
How would we write a regular expression to match a telephone number with the form (###) ###-####?
Attached to literals, character classes, or groups, these allow a match to repeat some number of times.
| Quantifier | Description |
|---|---|
* |
Match 0 or more |
+ |
Match 1 or more |
? |
Match 0 or 1 |
{3} |
Match exactly 3 |
{3,} |
Match 3 or more |
{3,5} |
Match 3, 4 or 5 |
How would we improve our previous regular expression for matching a telephone number with the form (###) ###-####?
For this exercise, treat the first token as the first name and everything after the first space as the last name, including prefixes and compound names. Match vowels (a, e, i, o, u) without regard to case.
For the following vector of randomly generated names, write a regular expression that,
detects if the person’s first name starts with a vowel (a,e,i,o,u)
detects if the person’s last name starts with a vowel
detects if either the person’s first or last name starts with a vowel
detects if neither the person’s first nor last name starts with a vowel
x = c("Jeremy Cruz", "Nathaniel Le", "Jasmine Chu", "Bradley Calderon Raygoza",
"Quinten Weller", "Katelien Kanamu-Hauanio", "Zuhriyaa al-Amen",
"Travale York", "Alexis Ahmed", "David Alcocer", "Jairo Martinez",
"Dwone Gallegos", "Amanda Sherwood", "Hadiyya el-Eid", "Shaimaaa al-Can",
"Sarah Love", "Shelby Villano", "Sundus al-Hashmi", "Dyani Loving",
"Shanelle Douglas")str_extract() returns the first match in each string, or NA when there is none. str_extract_all() returns every match, as a list with one element for each string.
At a given starting position, a greedy quantifier tries to consume as much text as possible while allowing the rest of the pattern to match. Adding ? makes it non-greedy, so it consumes as little as possible while still allowing the rest to match.
Groups allow you to connect pieces of a regular expression for modification or capture.
| Group | Description |
|---|---|
(a|b) |
match literal “a” or “b”, group either |
a(bc)? |
match “a” or “abc”, group bc or nothing |
(abc)def(ghi) |
match “abcdefghi”, group abc and ghi |
(?:abc) |
match “abc”, non-capturing group |
stringrstr_extract() returns the full match by default. Its group argument can select a capturing group. str_match() returns a matrix with the full match in the first column, followed by one column per capturing group.
str_match_all() returns a list with one matrix per input string. Each row is a match. The first column contains the full match, followed by one column per capturing group.
Both matches are retained, with the letter and digits captured separately. An input with no matches gets a matrix with zero rows.
Backreferences refer to previously captured groups using \1, \2, etc. They can be used in a replacement string, or within the pattern itself to match repeated text.
Flags change how a pattern is interpreted. In stringr they are set by wrapping the pattern in regex().
| Argument | Behavior |
|---|---|
ignore_case = TRUE |
match without regard to case |
multiline = TRUE |
^ and $ match at the start and end of each line |
dotall = TRUE |
. also matches \n |
comments = TRUE |
ignore whitespace and # comments in the pattern |
stringr’s functions are vectorized, so they can be used directly on the columns of a data frame.
Write a regular expression that will extract all of the phone numbers contained in the vector above. Expect 0, 1, 1, and 2 matches across the four inputs.
Once that works, use capturing groups with str_match_all() to retain every phone number while separating the area code from the remaining number. Keep the matches associated with their original input string.
re module| Function | Description |
|---|---|
re.search(p, s) |
Find the first match anywhere in the string, returns a Match or None. |
re.match(p, s) |
Match only at the start of the string. |
re.fullmatch(p, s) |
Match only if the entire string matches. |
re.findall(p, s) |
Return all matches or captured groups as a list of strings or tuples. |
re.finditer(p, s) |
Return all matches as an iterator of Match objects. |
re.sub(p, repl, s) |
Replace all matches with repl. |
re.split(p, s) |
Split the string at each match. |
re.compile(p) |
Compile a pattern into an object that can be reused. |
Most patterns shown so far translate directly, but regex engines differ in some details as well as their interfaces.
For example re does not support the POSIX classes like [[:digit:]]. Use \d, \s, \w, or an explicit class such as [a-z].
Python’s raw strings are written r"...", and the convention is to write every regex pattern as one.
re.search() takes a single string, so the vectorized str_detect() and str_subset() become comprehensions.
re.search() returns a Match object, or None when there is no match. None is falsy, so if m: is the usual test. Positions are 0-based and the end is exclusive.
re.sub() plays the role of both str_replace() and str_replace_all(). Backreferences in the replacement work the same way as in stringr.
A Match object gives access to the capture groups, with group(0) being the full match.
Flags are passed as an additional argument and combined with |. re.compile() stores a pattern and its flags in an object with the same methods as the re module.
| Behavior | stringr::regex() |
Python re |
|---|---|---|
| ignore case | ignore_case = TRUE |
re.IGNORECASE, re.I |
^ and $ match each line |
multiline = TRUE |
re.MULTILINE, re.M |
. matches \n |
dotall = TRUE |
re.DOTALL, re.S |
| whitespace and comments | comments = TRUE |
re.VERBOSE, re.X |
pandas and polarsMany .str methods in pandas and polars accept regular expressions. polars uses Rust’s regex engine, which does not support backreferences in patterns.
| Task | stringr |
Python |
|---|---|---|
| pattern string | "\\d+", r"(\d+)" |
r"\d+" |
| detect | str_detect() |
re.search() |
| subset | str_subset() |
comprehension with re.search() |
| first match | str_extract() |
m = re.search(p, s)m.group() if m else None |
| all matches | str_extract_all() |
re.findall(), re.finditer() |
| capture groups | str_match(), str_match_all() |
m.group(1), m.groups(), re.findall() |
| replace first | str_replace() |
re.sub(count=1) |
| replace all | str_replace_all() |
re.sub() |
| split | str_split() |
re.split() |
| no match: detect | str_detect() → FALSE |
re.search() → None |
| no match: first | str_extract() → NA |
check for None before .group() |
| no match: all | str_extract_all() → character(0) per input |
re.findall() → [] |
| literal pattern | fixed() |
in, s.replace(), re.escape() |
| flags | regex(ignore_case = TRUE) |
re.IGNORECASE |
| data frame columns | vectorized | .str.contains(), .str.extract() |
stringr gives R a consistent, vectorized set of string functions. Python’s equivalents are methods on a single str, applied to many strings with a comprehension or the .str accessor of pandas and polars.
Regular expression syntax is largely shared between the languages. stringr generally takes the string first and works over vectors. Python’s re takes the pattern first and works on one string. Return values depend on the operation.
Regex escapes sit on top of string escapes. Use raw strings, r"(...)" in R and r"..." in Python, to avoid doubling every backslash.
Build patterns up incrementally and test them against examples that should and should not match. If a pattern becomes unreadable, a few simpler steps are usually the better choice.
Sta 523 - Fall 2026