Information about a parsed document
Value
An object of class zuhtml_doc_info: a list with
type:"document"or"fragment";context: a fragment's context element,NAfor a document;nodes,attributes: counts;native_bytes: memory the document owns, outside R's heap;parse_peak_bytes: the most memory the parser held at once;input_bytes: size of the decoded UTF-8 input;encoding: the encoding the input was decoded from;encoding_source: where that encoding came from:"bom"(a byte-order mark),"argument"(theencodingargument),"meta"(a<meta>declaration),"default"(none of these, so UTF-8), or"string"for character input, which is already decoded;base_url: thebase_urlgiven tohtml_parse(), orNA;quirks_mode:"no-quirks","quirks"or"limited-quirks", as the doctype selected;problems: the number of parse problems kept, andproblems_truncated: whether more occurred thanmax_errors.
See also
Other parsing:
html_fragment(),
html_limits(),
html_parse(),
html_problems(),
zuhtml_info()
Examples
html_info(html_parse("<!DOCTYPE html><title>t</title><p>Hello"))
#> <zuhtml_doc_info>
#> type document
#> context NA
#> nodes 9
#> attributes 0
#> native_bytes 579
#> parse_peak_bytes 2,307
#> input_bytes 39
#> encoding UTF-8
#> encoding_source string
#> base_url NA
#> quirks_mode no-quirks
#> problems 0
#> problems_truncated FALSE