Runs the query behind a table created by tbl_delta(), tbl_parquet(),
tbl_csv() or tbl_json() and returns a stream that yields the result
in batches of at most chunk_size rows. Batches are produced as they are
read, so the whole result never needs to fit in memory. Runs the same
pre-flight checks as collect.tbl_az().
Arguments
- x
A
tbl_azproduced bytbl_delta(),tbl_parquet(),tbl_csv()ortbl_json().- chunk_size
Positive whole number. The maximum number of rows in each batch. DuckDB allocates each batch for
chunk_sizerows up front, so a very large value reserves a lot of memory.
Details
Requires the nanoarrow package. Read batches with
stream$get_next(), which returns NULL once the stream is exhausted,
or pass the stream to arrow::as_record_batch_reader().
Reading the stream
Read the stream to the end before running any other query on the same connection. DuckDB ends an open stream as soon as the connection runs another query, and the stream then reports that it is exhausted without an error, so the remaining rows are lost silently.
To stop reading early, call stream$release(). It frees the query and
its buffers at once, rather than when R next collects garbage.
See also
collect_arrow() to read the whole result at once.
Examples
if (FALSE) { # \dontrun{
# Requires a live Azure account, credentials, and network access.
conn <- az_conn()
sales <- tbl_delta(conn, "abfss://container@account/path/sales")
stream <- stream_arrow(sales, chunk_size = 100000)
while (!is.null(batch <- stream$get_next())) {
message(batch$length, " rows")
}
} # }