Skip to contents

Parsing

From a string, raw bytes or a file to a document, under limits.

html_parse() html_read()
Parse HTML
html_fragment()
Parse an HTML fragment
html_limits()
Resource limits for parsing and extraction
html_problems()
Parse problems recorded for a document
html_info()
Information about a parsed document
zuhtml_info()
Report the zuhtml build configuration

Selecting and navigating

Node values

html_text_clean()
Cleaned text for extraction
html_text()
Text content of nodes
html_strings()
Text pieces of nodes
html_markdown()
Convert HTML to Markdown
html_attr() html_attrs() html_classes()
Attributes of elements
html_name() html_namespace() html_type()
Node names, namespaces and types
html_serialize()
Serialize nodes as HTML

Extracting structures

html_table() html_tables()
Extract HTML tables as data frames
html_table_cells()
Cells of an HTML table
html_list()
Extract an HTML list
html_links()
Links in a document
html_url()
Resolve URLs in attributes
html_forms()
Forms and their controls

Page metadata

html_title()
Document title
html_meta()
Meta tags
html_json_ld()
JSON-LD blocks
html_microdata()
Microdata items

Conditions

zuhtml-conditions
Conditions raised by zuhtml