Skip to contents

Information about a parsed document

Usage

html_info(x)

Arguments

x

A zuhtml_document, or a zuhtml_nodeset for the document that owns it.

Value

An object of class zuhtml_doc_info: a list with

  • type: "document" or "fragment";

  • context: a fragment's context element, NA for a document;

  • nodes, attributes: counts;

  • native_bytes: memory the document owns, outside R's heap;

  • parse_peak_bytes: the most memory the parser held at once;

  • input_bytes: size of the decoded UTF-8 input;

  • encoding: the encoding the input was decoded from;

  • encoding_source: where that encoding came from: "bom" (a byte-order mark), "argument" (the encoding argument), "meta" (a <meta> declaration), "default" (none of these, so UTF-8), or "string" for character input, which is already decoded;

  • base_url: the base_url given to html_parse(), or NA;

  • quirks_mode: "no-quirks", "quirks" or "limited-quirks", as the doctype selected;

  • problems: the number of parse problems kept, and problems_truncated: whether more occurred than max_errors.

Examples

html_info(html_parse("<!DOCTYPE html><title>t</title><p>Hello"))
#> <zuhtml_doc_info>
#>   type               document
#>   context            NA
#>   nodes              9
#>   attributes         0
#>   native_bytes       579
#>   parse_peak_bytes   2,307
#>   input_bytes        39
#>   encoding           UTF-8
#>   encoding_source    string
#>   base_url           NA
#>   quirks_mode        no-quirks
#>   problems           0
#>   problems_truncated FALSE