Skip to contents

Runs the query behind a table created by tbl_delta(), tbl_parquet(), tbl_csv() or tbl_json(), reads the whole result into memory, and returns it as an Arrow stream, without converting it to an R data frame. Runs the same pre-flight checks as collect.tbl_az().

Usage

collect_arrow(x)

Arguments

x

A tbl_az produced by tbl_delta(), tbl_parquet(), tbl_csv() or tbl_json().

Value

A nanoarrow_array_stream whose batches are already in memory.

Details

The query has finished when collect_arrow() returns, so the connection is free for other queries. Use stream_arrow() instead to read a result that is too large to hold in memory.

Requires the nanoarrow package.

Reading the result

Like any Arrow stream, the result can be read only once. Convert it straight away, and keep the converted object:

Converting the result

as.data.frame() and tibble::as_tibble() convert the result with nanoarrow. The column types mostly match collect.tbl_az(), with these exceptions:

  • INTERVAL columns cannot be converted. Cast them in the query, or convert through arrow::as_arrow_table().

  • TIME columns become hms rather than difftime.

  • BLOB columns become blob rather than a plain list.

See also

stream_arrow() to read the result in batches.

Examples

if (FALSE) { # \dontrun{
# Requires a live Azure account, credentials, and network access.
conn <- az_conn()
sales <- tbl_delta(conn, "abfss://container@account/path/sales")
tab <- arrow::as_arrow_table(collect_arrow(sales))
} # }