html_attr() reads one attribute from every node; html_attrs() reads
all of them; html_classes() splits the class attribute into tokens.
Values are decoded (& becomes &). An attribute that is present
but empty is "", distinct from an absent one: test a boolean attribute
such as disabled by presence, !is.na(html_attr(x, "disabled")).
Value
html_attr(): a character vector as long asx;NAfor missing nodes.html_attrs(): a list as long asxof named character vectors, in source order;NA_character_for missing nodes.html_classes(): a list as long asxof character vectors of class tokens;NA_character_for missing nodes.
Details
Attribute names on HTML elements match regardless of ASCII case, as in a
browser; on SVG and MathML elements they match exactly (viewBox).
Namespaced attributes in foreign content are named with their prefix, as
in "xlink:href". Where an element has the same attribute more than
once, the first wins, as the HTML parser decides.
See also
Other node values:
html_markdown(),
html_name(),
html_serialize(),
html_strings(),
html_text(),
html_text_clean()
Examples
doc <- html_parse(
"<a href='/x' class='btn primary' data-id=7>Go</a><input disabled>"
)
body <- html_children(html_root(doc))[2]
nodes <- html_children(body)
html_attr(nodes, "href")
#> [1] "/x" NA
html_attr(nodes, "href", default = "")
#> [1] "/x" ""
html_attrs(nodes)
#> [[1]]
#> href class data-id
#> "/x" "btn primary" "7"
#>
#> [[2]]
#> disabled
#> ""
#>
html_classes(nodes)
#> [[1]]
#> [1] "btn" "primary"
#>
#> [[2]]
#> character(0)
#>
!is.na(html_attr(nodes, "disabled"))
#> [1] FALSE TRUE