Runs the query behind a table created by tbl_delta(), tbl_parquet(),
tbl_csv() or tbl_json(), reads the whole result into memory, and
returns it as an Arrow stream, without converting it to an R data frame.
Runs the same pre-flight checks as collect.tbl_az().
Arguments
- x
A
tbl_azproduced bytbl_delta(),tbl_parquet(),tbl_csv()ortbl_json().
Details
The query has finished when collect_arrow() returns, so the connection
is free for other queries. Use stream_arrow() instead to read a result
that is too large to hold in memory.
Requires the nanoarrow package.
Reading the result
Like any Arrow stream, the result can be read only once. Convert it straight away, and keep the converted object:
as.data.frame()ortibble::as_tibble()for an R data frame.arrow::as_arrow_table()for an Arrow Table, which can be reused.
Converting the result
as.data.frame() and tibble::as_tibble() convert the result with
nanoarrow. The column types mostly match collect.tbl_az(), with
these exceptions:
INTERVALcolumns cannot be converted. Cast them in the query, or convert througharrow::as_arrow_table().TIMEcolumns becomehmsrather thandifftime.BLOBcolumns becomeblobrather than a plain list.
See also
stream_arrow() to read the result in batches.