HTML is the language web pages are written in. It describes a page as nested elements marked by tags, which the browser renders into what we see. Web scraping is the reverse process, recovering the underlying data from that markup.
<html><head><title>This is a title</title></head><body><p align="center">Hello world!</p><br/><div class="name" id="first">John</div><div class="name" id="last">Doe</div><div class="contact"><div class="home">555-555-1234</div><div class="home">555-555-2345</div><div class="work">555-555-9999</div><div class="fax">555-555-8888</div></div></body></html>
Anatomy of an element
An HTML document is built from elements. Most have an opening tag, optional attributes, content, and a closing tag. Void elements such as br and img have no content or closing tag.
<div class="name" id="first">John</div>
Part
Here
Notes
name
div
what kind of element this is (p, a, table, div, …)
attribute(s)
class="name"id="first"
name and value pairs, any number of them
content
John
text, other elements, or both
Two attributes matter more than the rest when scraping:
class - used to label a group of similar elements. Multiple elements can share a class, and one element can have multiple classes (class="name bold").
id - labels a single element and should be unique within a page.
The document tree
Elements nest inside one another, so a page is a tree. Selectors describe elements both by what they are and by where they sit in the tree.
html
├── head
│ └── title
└── body
├── p
├── br
├── div.name#first
├── div.name#last
└── div.contact
├── div.home
├── div.home
├── div.work
└── div.fax
html is the root, an ancestor of everything else
body is the parent of p, and p is a child of body
p, br and the first three divs are siblings, since they share a parent
div.home is a descendant of body but not a child of it
Beautiful Soup
Beautiful Soup is the most widely used Python package for parsing and navigating HTML. It is installed as beautifulsoup4 and imported as bs4.
BeautifulSoup(html, parser) - parse HTML from a string, returning the document as a BeautifulSoup object.
The document and every element in it are Tag objects, with methods and attributes for
selecting - .select(), .select_one() by CSS selector, .find_all(), .find() by tag name and attributes.
extracting - .get_text(), .name, .attrs, and tag["attr"] or tag.get("attr") for one attribute.
traversing and modifying - moving between parents, children and siblings, and adding, replacing or removing elements. We will not cover these.
from bs4 import BeautifulSoupsoup = BeautifulSoup(html, "html.parser")soup.title
<title>This is a title</title>
type(soup)
<class 'bs4.BeautifulSoup'>
type(soup.title)
<class 'bs4.element.Tag'>
type(soup.title.string)
<class 'bs4.element.NavigableString'>
Class
What it is
BeautifulSoup
the whole document, behaves like a Tag
Tag
an element, with a name, attributes and children
NavigableString
a run of text inside a tag, behaves like a str
Parsers
The second argument to BeautifulSoup() picks the parser. Real pages are often not well formed, and each parser repairs broken HTML differently. The result is the tree you will be selecting from, so pick one parser and use it consistently. prettify() shows the repaired tree as indented HTML.
<html>
<body>
<p>
One
</p>
<p>
Two
</p>
<li>
Three
</li>
</body>
</html>
Tags as attributes
The simplest way to reach an element is by its tag name, as an attribute of the document or of another Tag. This returns the first element with that name, or None if there is none.
soup.title
<title>This is a title</title>
soup.p
<p align="center">Hello world!</p>
soup.div
<div class="name" id="first">John</div>
soup.body.div
<div class="name" id="first">John</div>
print(soup.body.dv)
None
Selecting elements
select() returns all matches as a list, and select_one() returns the first matching Tag, or None. Both can be called on the whole document or on any Tag.
The extraction methods belong to a single Tag. With more than one element you will need to use a comprehension.
soup.select(".name").get_text()
AttributeError: ResultSet object has no attribute "get_text". You're probably treating a list of elements like a single element. Did you call find_all() when you meant to call find()?
[d.get_text() for d in soup.select(".name")]
['John', 'Doe']
[d["id"] for d in soup.select(".name")]
['first', 'last']
CSS selectors
CSS (Cascading Style Sheets) is the language used to style web pages. A stylesheet is a list of rules, and each rule starts with a selector that says which elements the styling applies to.
Scraping tools borrow just the selector half of the language. It is a compact, standardized way of describing a set of elements, and it is already understood by a wide variety of tools.
Basic selectors
The simplest selectors match on an element’s tag name, its class, or its id. These can be written together with no spaces to require that all of them match, or separated by commas to match any of them.
Combinators join two selectors and match on the relationship between elements in the tree. The element that is returned is always the last one in the selector.
Selector
Example
Description
A B
body div
Select all <div> elements anywhere inside <body> (descendant)
A > B
body > div
Select all <div> elements directly inside <body> (child)
A + B
.work + div
Select the <div> immediately after a .work element (adjacent sibling)
A ~ B
.home ~ div
Select all <div> siblings that come after a .home element
Pseudo-classes start with a : and select on an element’s position among its siblings or on a condition that the other selectors cannot express.
Selector
Example
Description
:first-child
li:first-child
Select <li> elements that are the first child of their parent
:last-child
li:last-child
Select <li> elements that are the last child of their parent
:nth-child(n)
tr:nth-child(2)
Select <tr> elements that are the second child of their parent
:nth-child(odd)
tr:nth-child(odd)
Select every other <tr>, also even and formulas like 3n+1
:not(A)
div:not(.ad)
Select <div> elements that do not match .ad
:has(A)
div:has(img)
Select <div> elements that contain an <img>
Not every tool supports every pseudo-class, so a selector that works in the browser may fail in a scraping library. For example, the newer details:open works in Beautiful Soup but is an error in rvest.
There are usually many selectors that return the same elements. The goal is one that keeps working even if the page changes.
Prefer classes and IDs that describe the content (.price, #reviews) over ones that describe position (div > div:nth-child(3)).
Be as specific as needed and no more. A long chain of tags breaks the first time the site adds a div.
Be wary of machine-generated class names like .css-1q2w3e, which tend to change with every update of the site.
Anchor on a stable container, then select within it (#results .title).
Always check how many elements came back. Too many or too few is the first sign that the selector is wrong.
Not everything needs to be purely selector based - you can also process via Python or R after extraction.
Missing elements
When nothing matches, select() returns an empty list and select_one() returns None, so chaining a method onto a missing element is an error.
soup.select("span")
[]
print(soup.select_one("span"))
None
soup.select_one("span").get_text()
AttributeError: 'NoneType' object has no attribute 'get_text'
Attributes behave similarly. tag["attr"] raises a KeyError if the attribute is missing, while tag.get("attr") returns None.
print(p.get("class"))
None
find() and find_all()
Beautiful Soup also has its own search methods that match on tag name and attribute values. find_all() returns a list and find() returns the first match (or None). Both can be chained, as each Tag is searched in the same way as the full document.
Where find_all() earns its place over select() is in what it accepts as a filter. The tag name and each attribute can be a string, a list of strings, a regular expression, or True to match any value, and string= searches the text instead of the tags.
soup.find_all(["p", "title"])
[<title>This is a title</title>, <p align="center">Hello world!</p>]
get_text() returns the raw text of a tag and all of its descendants, including the source whitespace. The separator and strip arguments control how the pieces are joined.
.string is the text of a tag that has exactly one string as its content. It works for simple elements such as p or div.home, but a tag that contains other elements has no single string, so .string is None and get_text() must be used instead.
soup.p.string
'Hello world!'
soup.select_one("div.home").string
'555-555-1234'
print(contact.string)
None
Reading from the web
BeautifulSoup() does not accept a URL, so the page is downloaded first with requests and the text of the response is then parsed.
raise_for_status() turns a failed request (404, 500, …) into an exception. Without it we would happily parse the error page.
SelectorGadget
SelectorGadget is a browser extension (or bookmarklet) that builds a CSS selector interactively, by clicking on elements you want and elements you do not. When it cannot untangle a page, the browser’s inspector (right-click, Inspect) shows the exact tags, classes and ids around any element.
Sites use the robots exclusion standard to publish crawler rules in robots.txt. Following these rules does not mean you are authorized to use the website to collect the data. Also consider the site’s terms and the privacy of the people represented in the data.
You can find examples at all of your favorite websites: Google, Facebook, etc.
These files are meant to be machine readable, and urllib.robotparser from the standard library reads them. Check the rules for the same user agent that you will send with your requests.
from urllib.robotparser import RobotFileParseruser_agent ="sta523-example"rp = RobotFileParser("https://www.google.com/robots.txt")rp.read()rp.can_fetch(user_agent, "https://www.google.com/search?q=duke")rp.can_fetch(user_agent, "https://www.google.com/maps")
Rate limiting
Making requests too quickly can overload a server, get your IP blocked, or violate a site’s terms of service. Adding delays between requests is essential for responsible scraping.
The simplest approach is time.sleep() between requests. This loop uses rp from the previous slide, so all of the URLs must be on that site.
import timeuser_agent ="sta523-example"urls = ["https://www.google.com/search?q=duke","https://www.google.com/maps"]for url in urls:if rp.can_fetch(user_agent, url):try: r = requests.get(url, headers={"User-Agent": user_agent}, timeout=10) r.raise_for_status()finally: time.sleep(5)
Graceful error handling
When scraping many pages, retrieval or parsing can fail. Wrapping the request in try / except lets us capture those errors and continue, returning None on failure.
def safe_read(url):try: r = requests.get(url, timeout=10) r.raise_for_status()return BeautifulSoup(r.text, "html.parser")except requests.RequestException:returnNonepages = [safe_read(url) for url in urls]
Keeping the exception alongside the result lets us inspect what went wrong afterwards.
def safe_read(url):try: r = requests.get(url, timeout=10) r.raise_for_status()return {"result": BeautifulSoup(r.text, "html.parser"), "error": None}except requests.RequestException as e:return {"result": None, "error": e}
Using these selectors construct a data frame containing these results. Some movies will be missing one or more of these, so use None for any missing fields.
rvest
rvest is a package in the tidyverse that makes processing and manipulation of HTML data straightforward. It provides high-level functions for interacting with HTML via the xml2 package.
read_html() - read HTML data from a URL or character string.
html_elements() / html_nodes() - select specified elements from the HTML document using CSS selectors (or XPath).
html_element() / html_node() - select the first match within each input element using CSS selectors (or XPath).
html_table() - parse an HTML table into a data frame.
html_text() / html_text2() - extract a tag’s text content.
html_elements() returns every match as a node set. Unlike Beautiful Soup, the extraction functions are vectorized over node sets, so there is no need for a loop.
html_element() returns a missing node when nothing matches and the extraction functions then return NA, while html_elements() returns an empty node set and the extraction functions return an empty vector.
page =read_html(html)page |>html_element("span")
{xml_missing}
<NA>
page |>html_element("span") |>html_text()
[1] NA
page |>html_elements("span")
{xml_nodeset (0)}
page |>html_elements("span") |>html_text()
character(0)
This is why html_element() is the right choice when extracting one field per item, as the output stays aligned with the inputs.
par =read_html("<p> First sentence. Same line.<br>New line. </p>")
par |>html_text()
[1] "\n First sentence.\n Same line.New line.\n "
par |>html_text() |>cat(sep="\n")
First sentence.
Same line.New line.
par |>html_text2()
[1] "First sentence. Same line.\nNew line."
par |>html_text2() |>cat(sep="\n")
First sentence. Same line.
New line.
Scraping with polite
The polite package handles robots.txt, delays, and caching for us while maintaining the three pillars of a polite session:
seek permission
take it slowly
never ask twice
bow() reads the site’s robots.txt and returns a session object, nod() moves to a new path within the site, and scrape() makes the request (this is equivalent to read_html() and returns a parsed HTML object).
polite uses the larger of the requested delay (5 seconds by default) and the crawl delay in robots.txt.
Graceful error handling
purrr::possibly() wraps a function so that errors return a default value instead of stopping, which is the try / except pattern from earlier packaged as a function.
results |> purrr::map("result") # successful results (NULL on failure)results |> purrr::map("error") # error objects (NULL on success)
Summary
Beautiful Soup and rvest
Task
Python
rvest
download a page
requests.get(url).text
read_html(url)
parse HTML
BeautifulSoup(html, "html.parser")
read_html(html)
select all matches
select(css), find_all()
html_elements(css)
first match
select_one(css), find()
html_element(css), per input
no match
[], None
empty node set / missing node
text
get_text(), get_text(" ", strip=True)
html_text(), html_text2()
tag name
.name
html_name()
all attributes
.attrs
html_attrs()
one attribute
tag["href"], tag.get("href")
html_attr("href")
many elements
comprehension
vectorized
robots.txt
urllib.robotparser
polite::bow()
rate limiting
time.sleep()
polite::scrape(), Sys.sleep()
error handling
try / except
possibly(), safely()
Takeaways
HTML is hierarchical, and scraping is the work of picking the elements we want out of that hierarchy and flattening them into a data frame.
CSS selectors do the picking in both languages. Tools like SelectorGadget and the browser’s inspector help find a selector that matches what we want and nothing else.
Beautiful Soup’s methods work on one Tag at a time and are combined with comprehensions, while rvest’s functions are vectorized over a set of elements.
find_all() covers the cases a selector cannot, such as matching with a regular expression or an arbitrary function.
Just because a page can be scraped does not mean it should be. Check robots.txt and the terms of service, slow down your requests, and expect some of them to fail.