Columns are returned in requested order. Row groups are always returned in their physical file order, even if their selector is not sorted.
Arguments
- x
A
qio_parquet_fileobject.- ...
Reserved for future use.
- columns
Character vector of complete column paths, or
NULLfor all columns. Paths are matched exactly asschema()andnames()report them, dot-separated for nested leaves, and never by leaf name alone: two leaves may share a name under different parents. An unknown path is an error, and so is a path matching more than one leaf.- row_groups
Integer vector of 1-based row-group IDs, or
NULLfor all row groups.- batch_size
Positive number of rows decoded at a time. It bounds the reader's scratch memory, not the size of the result (see Details);
walk_batches()additionally uses it as the size of each batch.- int64
How 64-bit integer columns reach R.
"double"(the default) is exact from-2^53through2^53and returnsNAoutside it."integer64"returns bit64::integer64, which needs the suggestedbit64package and covers the signed 64-bit range except its lowest value:bit64reserves-9223372036854775808as its ownNA, so a column storingINT64_MINreads asNAin either mode. Either way values that cannot be represented becomeNA, and one warning naming the column is emitted for each column that lost values – once per column, however many values, row groups, or batches were affected, and never for a column that lost nothing. Unsigned 64-bit columns are never returned as negative numbers.- time
How
TIMEcolumns reach R:"numeric"(the default) returns seconds since midnight,"hms"returns hms::hms and needs the suggestedhmspackage. Neither returnsPOSIXct, because a time of day is not an instant.- tz
Time zone name,
"UTC"by default. A UTC-adjustedTIMESTAMPis an instant, sotzchanges only how it prints. A non-UTCTIMESTAMPis a wall clock with no zone stored, so its civil components are interpreted intz; base R decides ambiguous and nonexistent times at daylight-saving boundaries. The machine's local zone is never used implicitly.- verbose
Report what the read is about to do before doing it: the rows, columns, and row groups selected, the batch size, and the resolved
read_plan()for the selected columns only – not for the whole file, so it answers "what am I about to get". Written withmessage(), so it goes to stderr andsuppressMessages()silences it.
Details
collect() returns the whole selection, so batch_size does not bound the
result; use walk_batches() for that. It does bound the scratch memory the
reader allocates while decoding string and binary columns, which would
otherwise scale with the largest selected row group rather than with
anything the caller controls. Smaller batches lower peak memory and cost a
little throughput on dictionary-encoded text.
Nested and repeated columns are not materialized in qio 0.1.0. When a selection includes them, they are omitted and one message reports how many physical leaf columns were skipped. Nested reading is deferred to qio 0.2.0. If every selected column is nested, the result is a zero-column data frame with the selected number of rows.
See also
open_parquet() for the handle and for mmap and threads,
walk_batches() to process a file that does not fit in memory,
read_plan() to preview the R type of every column, and read_parquet(),
which is open_parquet() plus collect() for a whole file.
Examples
path <- tempfile(fileext = ".parquet")
write_parquet(mtcars, path)
pf <- open_parquet(path)
collect(pf, columns = c("mpg", "cyl"))
#> mpg cyl
#> 1 21.0 6
#> 2 21.0 6
#> 3 22.8 4
#> 4 21.4 6
#> 5 18.7 8
#> 6 18.1 6
#> 7 14.3 8
#> 8 24.4 4
#> 9 22.8 4
#> 10 19.2 6
#> 11 17.8 6
#> 12 16.4 8
#> 13 17.3 8
#> 14 15.2 8
#> 15 10.4 8
#> 16 10.4 8
#> 17 14.7 8
#> 18 32.4 4
#> 19 30.4 4
#> 20 33.9 4
#> 21 21.5 4
#> 22 15.5 8
#> 23 15.2 8
#> 24 13.3 8
#> 25 19.2 8
#> 26 27.3 4
#> 27 26.0 4
#> 28 30.4 4
#> 29 15.8 8
#> 30 19.7 6
#> 31 15.0 8
#> 32 21.4 4
close_parquet(pf)