Skip to contents

qio reads and writes Apache Parquet files from R. Read a whole file with read_parquet(), or open a larger one with open_parquet() to inspect its schema and read only the columns, row groups, or batches you need. It is built on the bundled C library carquet and has no required R package dependencies.

Installation

qio is not on CRAN yet. Install the development version from GitHub:

pak::pak("pedrobtz/qio")

Building from source needs GNU make and a C compiler. Zstandard and LZ4 are bundled; zlib comes from the system, or from Rtools on Windows.

Usage

write_parquet() and read_parquet() handle a whole file at a time.

write_parquet(mtcars, "mtcars.parquet")

cars <- read_parquet("mtcars.parquet")

Read part of a file by naming columns or row groups. A column you do not select is never decompressed, which is the cheapest speed-up available on a wide file.

read_parquet("mtcars.parquet", columns = c("mpg", "cyl"))

open_parquet() returns a handle. Inspecting one is cheap because it reads the footer, not the data, so you can look before deciding what to read.

# several row groups, so there is something to select
write_parquet(mtcars, "mtcars.parquet", row_group_size = 16)

pf <- open_parquet("mtcars.parquet")
pf
#> <qio_parquet_file>
#> mtcars.parquet
#> 32 rows x 11 columns; 2 row groups

schema(pf)      # column paths, Parquet types, nullability
row_groups(pf)  # rows and bytes per group
read_plan(pf)   # the R type each column will become

collect(pf, columns = c("mpg", "cyl"))
collect(pf, row_groups = 1)

close_parquet(pf)

Use walk_batches() for a file that does not fit in memory. It calls your function once per batch and keeps only one batch alive at a time.

pf <- open_parquet("big.parquet")

walk_batches(pf, batch_size = 100000, FUN = function(batch, index) {
  # one data frame at a time
})

close_parquet(pf)

Handles hold an open file, so close them when you are done. read_parquet() opens and closes one for you.

Type mapping

Parquet describes storage with a physical type and meaning with an optional logical type. qio reads them like this by default:

Parquet Logical type R
BOOLEAN logical
INT32 integer
INT32 DATE Date
INT32, INT64 TIME double (seconds)
INT64 double
INT64 TIMESTAMP POSIXct (UTC)
INT96 POSIXct (UTC)
FLOAT, DOUBLE double
BYTE_ARRAY STRING, ENUM, JSON character
BYTE_ARRAY list of raw
FIXED_LEN_BYTE_ARRAY list of raw
FIXED_LEN_BYTE_ARRAY UUID character
FIXED_LEN_BYTE_ARRAY FLOAT16 double
any DECIMAL double

Bytes are text only when the file says so, which is why an unannotated BYTE_ARRAY stays raw. INT64 is exact through 2^53 and NA beyond it; pass int64 = "integer64" for the full range. Nulls become NA. Nested and repeated columns are skipped for now.

Writing infers the reverse, and parquet_schema() overrides it per column:

R Parquet
logical BOOLEAN
integer INT32
double DOUBLE
character, factor BYTE_ARRAY + STRING
Date INT32 + DATE
POSIXct INT64 + TIMESTAMP (UTC, microseconds)
types <- parquet_schema(mpg = "FLOAT", cyl = "INT64")
write_parquet(mtcars, "mtcars.parquet", schema = types)

read_plan() reports what any file will produce before you read it, and ?qio-types documents every mapping, including where precision is lost.

Other functions

write_parquet() takes compression, row_group_size, metadata, sorted_by, and append. Beyond schema() and row_groups(), a handle can report column_chunks(), column_statistics(), page_index(), metadata(), and bloom_filter_may_contain(). validate_parquet() checks that a file is structurally sound without reading it.

See ?qio-limitations for what qio deliberately does not do.

Licensing

qio is MIT licensed. It bundles third-party C sources, each under its own license and shipped with its license file:

Bundled License Location
carquet MIT src/carquet/LICENSE
Snappy (in carquet) BSD-3-Clause src/carquet/compression/snappy.c
Zstandard BSD-3-Clause src/zstd/LICENSE
LZ4 BSD-2-Clause src/lz4/LICENSE

inst/COPYRIGHTS lists every copyright holder, the files each covers, and the modifications qio makes; DESCRIPTION points at it through its Copyright field. Each bundled library is pinned to an exact upstream commit.

qio carries local patches to carquet; most fix defects that silently corrupted or rejected valid data. They live as individual commits on the qio branch of a carquet fork, which is what src/carquet is vendored from, so each one can be read on its own and offered upstream. The reason for each is recorded in the repository’s .agents/VENDORED.md, which is not shipped in the source package.