Reads an Apache Parquet file into a data frame.
Arguments
- file
Path to a Parquet file, or an
http://,https://,ftp://,ftps://orfile://URL. A URL is downloaded to the session temporary directory in full before any of it is read, and the copy is removed when the read finishes; see qio-limitations.- ...
Must be empty. Every argument after it is name-only, matching
collect(),walk_batches()andread_plan(), which take the same arguments the same way.- columns
Character vector of complete column paths, or
NULL(the default) for all columns; seecollect().- row_groups
Integer vector of 1-based row-group IDs, or
NULL(the default) for all row groups.- int64
How 64-bit integer columns reach R; see
collect().- time
How
TIMEcolumns reach R; seecollect().- tz
Time zone for
TIMESTAMPcolumns; seecollect().- verbose
Report the read plan before reading; see
collect().
Details
Column types are mapped from Parquet as follows: BOOLEAN to logical,
INT32 to integer, and INT64/FLOAT/DOUBLE to double. A BYTE_ARRAY
becomes character only when the file annotates it STRING, ENUM, or
JSON; without an annotation it is arbitrary bytes and is returned as a
list of raw vectors, because assuming UTF-8 the file never claimed would
corrupt binary data. An INT32 column annotated DATE is returned
as a Date, a UTC-adjusted TIMESTAMP (physical INT64) as a POSIXct in
UTC, and a legacy INT96 timestamp as a POSIXct in UTC (interpreting its
Julian-day and nanosecond-of-day parts as an instant). Parquet nulls become
NA. INT64 values are returned as doubles and lose precision beyond 2^53,
which for microsecond and nanosecond timestamps can drop sub-second precision
far from the epoch. Use read_plan() to preview the R type of each column
before reading.
Nested and repeated columns are skipped with one message. Nested reading is deferred to qio 0.2.0.
The file is memory-mapped for the duration of the read (falling back to buffered reads if mapping fails) so columns decode in parallel; the mapping is released before the function returns.
columns and row_groups read part of a file and are passed straight to
collect(). Selecting columns is the single largest speedup available on a
wide file, because a column that is not selected is never decompressed.
Use open_parquet() with collect() for the rest: batch_size, mmap,
threads, and verify_checksums.
See also
write_parquet(), open_parquet() and collect() to read part of
a file, walk_batches() for a file larger than memory, and read_plan()
to preview the R type of every column before reading.
Examples
path <- tempfile(fileext = ".parquet")
write_parquet(data.frame(x = 1:3, y = c("a", "b", NA)), path)
read_parquet(path)
#> x y
#> 1 1 a
#> 2 2 b
#> 3 3 <NA>