Changelog
Source:NEWS.md
zuhtml 0.1.0
First release.
Bundles the ‘Gumbo’ HTML5 parser 0.14.0 from the maintained fork at https://codeberg.org/gumbo-parser/gumbo-parser, with local patches that add a parse-time nesting-depth limit, remove the library’s only
printf(), and fix two memory-safety bugs in its<selectedcontent>support that fuzzing found (a use-after-free and a NULL dereference, both reachable from untrusted HTML) and an uninitialized read in fragment parsing. No system library is needed.New
html_parse()andhtml_read()parse HTML from a string, raw vector or local file under explicit resource limits from the newhtml_limits(). Nesting depth is bounded while parsing, and all parser memory goes through an allocation ledger with a budget, so a failed parse always releases everything. Raw input is decoded from a byte-order mark, an explicit encoding, the page’s<meta>declaration or UTF-8; invalid input is an error, never silently replaced.Parsed documents are converted into a compact immutable tree that holds no reference to the parser or the input. All 1,686 applicable html5lib tree-construction tests that the bundled parser ships produce the expected tree.
New
html_fragment()parses markup in the context of a given element, asinnerHTMLdoes.Nodes are
zuhtml_nodesets: vectors of nodes tied to their document, with missing nodes where an aligned operation has no answer. Navigate withhtml_children(),html_parent(),html_ancestors(),html_next_sibling(),html_previous_sibling(),html_root(),html_document()andhtml_template_content(); read values withhtml_name(),html_namespace(),html_type(),html_attr(),html_attrs(),html_classes()andhtml_text().html_info()describes a document.lapply()and friends over a nodeset pass one node at a time;rep(),rev(),unique()andc()keep nodesets nodesets.New extraction functions.
html_text_clean()gives text as a reader wants it: scripts and styles skipped, whitespace collapsed outside<pre>, line breaks at<br>and block elements.html_list()reads a<ul>/<ol>as text or a tree, without nested items leaking into their parents.html_table()andhtml_tables()read tables into data frames of character columns, with row and column spans,rowspan="0", row groups, header detection and an error rather than a silent overwrite for overlapping cells.html_url()resolves URL attributes with RFC 3986 reference resolution, honouring<base href>, andhtml_links()lists a page’s links.New
html_elements(),html_element(),html_matches()andhtml_filter()select elements with a documented subset of CSS selectors: type, universal, ID, class and attribute selectors (with theiandsflags), the four combinators, selector lists,:scope,:root,:empty,:first-child,:last-child,:only-child,:nth-child(),:nth-of-type()and:not(). Anything else is azuhtml_selector_errorpointing at the offending position.html_element()keeps one result per input node, so extracted columns stay aligned.New
html_closest()finds each node’s nearest ancestor matching a selector, andhtml_strings()returns the text pieceshtml_text()joins, keeping element boundaries.html_serialize(pretty = TRUE)lays block-level elements out on indented lines for reading.New
html_title(),html_meta(),html_json_ld()andhtml_microdata()read page metadata: the document title, every<meta>tag (OpenGraph, Twitter cards and Dublin Core included), JSON-LD blocks (parsed with jsonlite if asked) and microdata items per the HTML standard,itemrefincluded.New
html_table_cells()returns one row per table cell with its grid position, spans, section, whether it is a header, and its links.html_tables()gainsmatch =to keep tables whose text matches a regular expression.html_table()gainsconvert =,decimal =andthousands =to convert columns that are entirely numbers or logicals, keeping identifiers with leading zeros as text.New
html_markdown()converts nodes to CommonMark: headings, paragraphs, emphasis, code, fenced code blocks, block quotes, nested lists, links and images with resolved URLs, and GFM pipe tables for data tables. Markdown-significant characters in text are escaped, and only emphasis that CommonMark parses back is written.The
<meta>declaration of raw input is found as browsers find it, by the HTML standard’s prescan of the first 1024 bytes, and read with the Encoding Standard’s labels (soiso-8859-1means windows-1252).html_info()reports where the encoding came from inencoding_source.New
html_forms()describes each form and the controls it owns, by the HTML standard’s form-owner rules (form=references included), with DOM values for every control type and the options of each select. It only inspects: nothing is submitted.html_read()reads a URL (through base R’surl(), which also sets the defaultbase_url) or any connection, such asgzfile(), as well as a file path. Reading stops once the input passes four timesmax_input.New
html_serialize()(andas.character()on nodesets) writes nodes as normalized HTML following the WHATWG serialization algorithm, outer or inner.New
html_problems()lists the parse errors the parser repaired, with package-owned codes and positions.Errors are classed conditions under
zuhtml_error; see?zuhtml-conditions.New
zuhtml_info()reports the bundled parser version and patches, and self-tests the compiled parser.