html_table() reads one <table> element into a data frame of
character columns. html_tables() finds tables and reads each.
Usage
html_table(
x,
header = "auto",
trim = TRUE,
na = character(),
convert = FALSE,
decimal = ".",
thousands = NULL,
limits = html_limits()
)
html_tables(x, css = "table", match = NULL, ...)Arguments
- x
For
html_table(), a nodeset holding exactly one<table>element; forhtml_tables(), azuhtml_documentorzuhtml_nodesetto search.- header
How to find header rows; see the Headers section.
- trim
Whether to trim each cell's text; see
html_text_clean().- na
Strings that become
NAafter trimming, such asc("", "-"). By default no text is treated as missing.- convert
If
TRUE, convert columns that are entirely numbers or logicals; see the Conversion section.- decimal
The decimal mark for
convert: a single character.- thousands
The grouping mark for
convert: a single character other thandecimal, orNULLfor none.- limits
Resource limits from
html_limits();max_table_cellsbounds the grid, spans included.- css
For
html_tables(), the selector of tables to read. Only the outermost matching tables are read; nested ones remain part of their parents' cells.- match
For
html_tables(), a regular expression (grepl()) that a table's cleaned text must match for the table to be read, orNULLto read every table.- ...
For
html_tables(), arguments passed on tohtml_table().
Value
html_table(): a data frame of character columns. A table with
no rows gives a data frame with no rows and no columns; a table with
only header rows gives no rows and the named columns.
html_tables(): a list of such data frames, list() when there are
none.
The grid
Rows are read in logical order – every <thead>, then the <tbody>
sections (and any rows directly in the table), then every <tfoot> –
and footer rows are ordinary data rows. Rows of nested tables never
become rows of the outer table, and a nested table's text is left out of
the cell that holds it (it still separates the text around it with a
line break).
Each cell is placed at the next free column of its row. colspan and
rowspan are read with HTML's rules for non-negative integers: a
missing or invalid value is 1, colspan="0" is 1, and values are capped
at 1000 columns and 65534 rows. rowspan="0" extends to the end of the
cell's row group, and a larger rowspan is cut silently at the end of
its group. A spanned value is repeated in every slot it covers. A cell
whose span would cover a slot another cell already covers is an error of
class zuhtml_table_structure_error, never a silent overwrite. Slots no
cell covers, in ragged rows, are NA; an empty cell is "".
Headers
"auto": the rows of<thead>sections if there are any, otherwise the leading rows whose cells are all<th>(a<th>among<td>s does not make a header row);TRUE: the first row;FALSE: none;a vector of row numbers: those rows.
Header rows are removed from the data. A column's name joins the
non-empty header texts above it with " / ", merging repeats that a
colspan created. Blank names become V1, V2, ... by column
position, and duplicates are made unique with make.unique(). Names are
not otherwise made syntactic.
Cell text follows html_text_clean(). All columns are character unless
convert = TRUE.
Conversion
With convert = TRUE, each column is converted on its own, and only when
every non-missing value converts; otherwise it stays character, so
conversion never turns a value into NA. Values are compared after
removing surrounding whitespace.
TRUE,FALSE,true,false,TrueandFalsemake a logical column.Decimal numbers, with an optional sign and exponent, make an integer column when all are whole and fit, otherwise a double column.
decimalis the decimal mark.thousands, if given, is a grouping mark allowed between groups of three digits, as in1,234,567.5; a badly grouped value such as1,23keeps the column character.A value with a leading zero, such as
"0012", keeps the column character: it is an identifier, not a number."0"and"0.5"are numbers.Inf,NaN,NA, currency symbols and percentages are not numbers; list such strings inna, or convert them yourself. A column with no non-missing values stays character.
See also
Other extraction:
html_forms(),
html_links(),
html_list(),
html_table_cells(),
html_url()
Examples
doc <- html_parse(paste0(
"<table><thead><tr><th>Product<th>Price</thead>",
"<tbody><tr><td>0012<td>12.50<tr><td>0034<td>9.00</tbody></table>"
))
html_table(html_element(doc, "table"))
#> Product Price
#> 1 0012 12.50
#> 2 0034 9.00
doc <- html_parse(paste0(
"<table><tr><th rowspan=2>Region<th colspan=2>Sales",
"<tr><th>2025<th>2026<tr><td>North<td>1<td>2</table>"
))
html_tables(doc)
#> [[1]]
#> Region Sales / 2025 Sales / 2026
#> 1 North 1 2
#>
html_tables(doc, match = "North")
#> [[1]]
#> Region Sales / 2025 Sales / 2026
#> 1 North 1 2
#>
doc <- html_parse(paste0(
"<table><tr><th>Id<th>Amount<th>Paid",
"<tr><td>007<td>1.234,50<td>true<tr><td>012<td>99<td>false</table>"
))
str(html_table(html_element(doc, "table"), convert = TRUE,
decimal = ",", thousands = "."))
#> 'data.frame': 2 obs. of 3 variables:
#> $ Id : chr "007" "012"
#> $ Amount: num 1234 99
#> $ Paid : logi TRUE FALSE