zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.
- Malformed markup is repaired by the HTML parsing algorithm, not rejected.
html_problems()lists what was repaired. - Select elements with a documented subset of CSS. Anything outside the subset is an error, never a partial match.
- Extract text, attributes, links, lists and tables. Tables handle row and column spans and keep every column as character unless you ask for conversion, so
"0012"stays"0012". - Read page metadata (title,
<meta>tags, JSON-LD, microdata) and forms, and convert any node to Markdown withhtml_markdown(). - Raw bytes are decoded as a browser decodes them: from a byte-order mark, your
encoding, or the page’s<meta charset>. - Every call runs under explicit limits on input size, native memory and nesting depth, and every error is a classed condition.
html_read() reads a file, a URL or any R connection. zuhtml has no HTTP client of its own: a URL goes through base R’s url(), and a fetcher that needs headers or authentication hands it the body. zuhtml does not run JavaScript or sanitize HTML.
Installation
install.packages("zuhtml")The development version, from GitHub:
# install.packages("pak")
pak::pak("pedrobtz/zuhtml")Example
library(zuhtml)
doc <- html_parse('
<div class=product><h2>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a></div>
<div class=product><h2>Genmaicha</h2>
<a href="genmaicha.html">details</a></div>',
base_url = "https://example.org/shop/"
)
cards <- html_elements(doc, ".product")
data.frame(
name = html_text_clean(html_element(cards, "h2")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.htmlhtml_element() returns one result per card, with a missing value where a card has no price, so the columns stay aligned.
html_read() reads a file, a connection or a URL. A URL becomes the document’s base URL, so relative links resolve against the page:
doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"
rows <- html_elements(doc, "table tr")
views <- data.frame(
topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
url = html_url(html_element(rows, "a"))
)
head(views, 3)
#> topic url
#> 1 Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science https://cran.r-project.org/web/views/Agriculture.html
#> 3 Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.htmlThe getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.
Licence
zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and LICENSE.note explains how the two apply.