Builds a read plan from a Parquet schema: one row per physical leaf column describing the R type each column will materialize as, whether it can be collected, and why not when it cannot. The plan is a pure function of the schema, so it is cheap to compute and inspect before reading any data.
Usage
read_plan(x, ...)
# S3 method for class 'qio_parquet_file'
read_plan(
x,
...,
int64 = c("double", "integer64"),
time = c("numeric", "hms"),
tz = "UTC"
)
# S3 method for class 'character'
read_plan(
x,
...,
int64 = c("double", "integer64"),
time = c("numeric", "hms"),
tz = "UTC"
)
# S3 method for class 'data.frame'
read_plan(
x,
...,
int64 = c("double", "integer64"),
time = c("numeric", "hms"),
tz = "UTC"
)Arguments
- x
A Parquet file path, a
qio_parquet_fileobject, or the data frame returned byschema().- ...
Reserved for future use.
- int64
How 64-bit integer columns reach R; see
collect(). The plan reports the resultingr_typeandconverter, so it can be inspected for exactly the read that will follow.- time
How
TIMEcolumns reach R; seecollect().- tz
Time zone for
TIMESTAMPcolumns; seecollect().
Value
A qio_read_plan data frame with one row per physical leaf column
and the columns:
column1-based physical column index.
name,pathColumn name and dotted path.
physical_type,logical_typeParquet physical type and logical annotation (
NAwhen absent).r_typeTarget R type, or
NAwhen the column cannot be collected.converterStable identifier of the conversion the reader uses.
nullableWhether the column can contain nulls.
nestedWhether the leaf belongs to a nested or repeated field.
collectibleWhether
collect()can currently materialize the column.noteReason a column is not collectible, or a pending logical annotation;
NAotherwise.
Details
The plan reflects what collect(), read_parquet(), and walk_batches()
actually do today. A logical annotation overrides the physical fallback in
parquet_type_mapping(), and the annotations qio resolves are:
DATEtoDate, and legacy physicalINT96toPOSIXct.TIMESTAMPtoPOSIXct, UTC-adjusted or interpreted intz.TIMEto seconds or hms::hms, selected bytime.INTEGERat any width and sign, including unsigned 64-bit, which withINT64is selected byint64.STRING,ENUM, andJSONto character; everything else stored as bytes stays a list of raw vectors.UUIDto canonical text andFLOAT16to double.DECIMALto double with the scale applied, from either integer or binary storage.
Unimplemented annotations are reported in note and retain their physical
fallback type. converter names the exact conversion the reader will run,
so it distinguishes cases that share an r_type.
A path is enough – read_plan() opens the file, reads the footer, and
closes it again, so no handle is needed to inspect a file before reading it.
Passing an open open_parquet() handle, or the data frame from schema(),
produces exactly the same plan; use those when a handle is already open or
when the schema has already been fetched.
Pass the same int64, time, and tz the read will use. The plan resolves
them, so r_type and converter describe that read rather than a default
one.
Examples
path <- tempfile(fileext = ".parquet")
write_parquet(data.frame(x = 1:3, y = c("a", "b", NA)), path)
# A path is enough; no handle is needed.
read_plan(path)
#> <qio_read_plan: 2 columns, 2 collectible>
#> column name path physical_type logical_type r_type converter nullable
#> 1 1 x x INT32 <NA> integer int32 FALSE
#> 2 2 y y BYTE_ARRAY STRING character text TRUE
#> nested collectible note
#> 1 FALSE TRUE <NA>
#> 2 FALSE TRUE <NA>
# The plan answers for the read you are about to do, not a default one.
read_plan(path, int64 = "integer64")
#> <qio_read_plan: 2 columns, 2 collectible>
#> column name path physical_type logical_type r_type converter nullable
#> 1 1 x x INT32 <NA> integer int32 FALSE
#> 2 2 y y BYTE_ARRAY STRING character text TRUE
#> nested collectible note
#> 1 FALSE TRUE <NA>
#> 2 FALSE TRUE <NA>
# An open handle and a schema() data frame give the same plan.
pf <- open_parquet(path)
identical(read_plan(pf), read_plan(path))
#> [1] TRUE
close_parquet(pf)