Reports the per-column, per-row-group statistics recorded in the file: value and null counts, and the minimum and maximum bounds.
Value
A data frame with one row per column chunk, ordered by row group and then by column, with the columns:
row_group1-based row-group ID, as
row_groups()reports it.column1-based physical column index, as
schema()reports it.pathComplete dotted column path; see
schema().num_valuesValues the statistics cover, nulls included.
null_countNulls the writer recorded, or
NAwhen it recorded none.NAmeans unknown, not zero.distinct_countDistinct values the writer recorded, or
NA. Most writers omit it, soNAis the common case.min,maxList columns of physical-level bounds, one element per row;
NULLwhen absent or undecodable. See Details.
Counts are doubles rather than integers, because a chunk can hold more
values than .Machine$integer.max.
Details
These are claims made by whoever wrote the file, not facts qio verifies. A reader that skips a row group on them is trusting that writer. qio does not use them to skip anything.
min and max are list columns, because one file can hold columns of
different types. Each element holds the bound decoded at the physical
level: an INT64 bound stays a number rather than becoming a POSIXct, and
a decimal is not scaled, since a bound is a sort key rather than a value to
compute with. Text columns are the exception and decode to character. An
element is NULL when the bound is absent, or is present but the wrong width
for its type.
Examples
path <- tempfile(fileext = ".parquet")
write_parquet(data.frame(n = 1:100), path, row_group_size = 25)
pf <- open_parquet(path)
stats <- column_statistics(pf)
stats[c("row_group", "path", "null_count")]
#> row_group path null_count
#> 1 1 n 0
#> 2 2 n 0
#> 3 3 n 0
#> 4 4 n 0
unlist(stats$min)
#> [1] 1 26 51 76
close_parquet(pf)