Finds the <a> and <area> elements with an href below the given
nodes (and the nodes themselves), in document order, without
duplicates. Duplicate destinations are kept, and so are fragment,
mailto: and tel: links.
Arguments
- x
A
zuhtml_documentorzuhtml_nodeset.- absolute
If
TRUE,urlis resolved withhtml_url(); otherwise it ishrefunchanged.
Value
A data frame with one row per link and character columns text
(the link's cleaned text, see html_text_clean()), href (the
attribute as written, decoded) and url.
See also
Other extraction:
html_forms(),
html_list(),
html_table(),
html_table_cells(),
html_url()
Examples
doc <- html_parse(
"<nav><a href='/'>Home</a> <a href='about.html'>About</a></nav>",
base_url = "https://example.org/site/"
)
html_links(doc)
#> text href url
#> 1 Home / /
#> 2 About about.html about.html
html_links(doc, absolute = TRUE)
#> text href url
#> 1 Home / https://example.org/
#> 2 About about.html https://example.org/site/about.html