Files
json/docs/mkdocs/docs/api/basic_json_document/read.md
Niels Lohmann 6ee5e89803 Scan json_view strings with SIMD and index large objects
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.

Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.

On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.

Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.

simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 10:48:07 +02:00

2.7 KiB

nlohmann::basic_json_document::read

template<typename InputType>
void read(InputType&& input,
         const bool allow_exceptions = true,
         const bool ignore_comments = false,
         const bool ignore_trailing_commas = false);

(Re-)parses input into #!cpp *this, discarding the document's previous value and reusing its memory (the node index, the decoded-string buffer, and, if applicable, the owned copy of the text) rather than allocating a fresh document. parse() is implemented in terms of this function, applied to a default-constructed document.

Template parameters

InputType
A compatible input; see parse.

Parameters

input (in)
Input to parse from.
allow_exceptions (in)
whether to throw exceptions in case of a parse error (optional, #!cpp true by default)
ignore_comments (in)
whether comments should be ignored and treated like whitespace (#!cpp true) or yield a parse error (#!cpp false); (optional, #!cpp false by default)
ignore_trailing_commas (in)
whether trailing commas in arrays or objects should be ignored and treated like whitespace (#!cpp true) or yield a parse error (#!cpp false); (optional, #!cpp false by default)

Exceptions

Same as parse.

Complexity

Linear in the length of the input.

Notes

Every view taken from #!cpp *this before the call -- including the previous root() -- is invalidated, whether or not the new parse succeeds; take fresh views from root() afterward.

input is borrowed or owned by the same rules as parse(); a document can borrow on one call and own on the next, since ownership is decided freshly each time.

Reusing a document matters most for large inputs: the operating system provides the memory of a fresh node index one page at a time, and every page costs a page fault the first time it is written. On x86-64 Linux (4 KiB pages), parsing a 55 MB document into a reused document took about 40 % less time than parsing it into a fresh one. Programs that parse many documents of similar size should therefore keep one document and call read().

Examples

??? example

The example below parses a sequence of messages into the same document, reusing its memory instead of allocating
a new document for each one.

```cpp
--8<-- "examples/basic_json_document__read.cpp"
```

Output:

```json
--8<-- "examples/basic_json_document__read.output"
```

See also

  • parse - deserialize from a compatible input
  • root - the view of the root value

Version history

  • Added in version 3.13.0.