NDJSON (application/x-ndjson, also called JSON Lines) is one JSON value per
line. json_parse_ndjson() turns such a body into a list of records;
json_write_ndjson() and json_write_ndjson_raw() go the other way.
Usage
json_parse_ndjson(x, simplify = TRUE, data_frame = FALSE)
json_write_ndjson(x, auto_unbox = TRUE)
json_write_ndjson_raw(x, auto_unbox = TRUE)Arguments
- x
For
json_parse_ndjson(), a single string or a raw vector of NDJSON bytes. For the writers, a list of records or a data frame — a data frame is written one object per row, which is the shape NDJSON exists for.- simplify
Passed through to
json_parse()for each record.- data_frame
Passed through to
json_parse()for each record. Note that records are parsed one at a time, so this turns an array inside a record into a data frame; it does not make a frame out of the stream.- auto_unbox
Passed through to
json_write()for each record.
Value
json_parse_ndjson() returns a list with one element per record,
each being what json_parse() returns for that line. json_write_ndjson()
returns a length-1 character vector and json_write_ndjson_raw() a raw
vector of UTF-8 bytes; both hold one record per line and end with a
trailing newline, so appending another record is always valid.
Details
Each line is parsed exactly as json_parse() would parse it, and each record
is written exactly as json_write() would write it, so the type mappings
documented there apply unchanged. The result of parsing is always a list,
one element per record, however uniform the records are — records are
independent documents, and an NDJSON body is not a table.
Blank lines are skipped rather than treated as records, and \r\n line
endings are accepted. A parse failure reports the line number, because
"invalid JSON at byte 41827" is not useful in a body of 10,000 records.
Why line framing is safe
A raw newline byte cannot appear inside a JSON string — it must be escaped as
\\n — and zujson's writer always escapes it. A serialized record therefore
never contains a bare newline, so splitting on newlines can never cut a
record in half. This is the property the format rests on.
It is also why there is no pretty argument: indented JSON contains
newlines, and a newline inside a record is exactly what NDJSON framing cannot
survive. Pretty-printed NDJSON is corrupt, not prettier.
See also
json_parse() and json_write() for single documents.
Examples
body <- '{"id":1,"ok":true}\n{"id":2,"ok":false}\n'
json_parse_ndjson(body)
#> [[1]]
#> [[1]]$id
#> [1] 1
#>
#> [[1]]$ok
#> [1] TRUE
#>
#>
#> [[2]]
#> [[2]]$id
#> [1] 2
#>
#> [[2]]$ok
#> [1] FALSE
#>
#>
# a data frame is one record per row
cat(json_write_ndjson(data.frame(id = 1:2, nm = c("a", "b"))))
#> {"id":1,"nm":"a"}
#> {"id":2,"nm":"b"}
# round trip
recs <- list(list(a = 1L), list(b = "x"))
identical(json_parse_ndjson(json_write_ndjson(recs)), recs)
#> [1] TRUE
# bytes, ready to be an HTTP request body
json_write_ndjson_raw(list(list(a = 1L), list(a = 2L)))
#> [1] 7b 22 61 22 3a 31 7d 0a 7b 22 61 22 3a 32 7d 0a