The node index section says what extra holds for objects and how large objects are indexed. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
20 KiB
Architecture
This page gives a high-level overview of the library's architecture. It should help new contributors to get an idea of the used concepts and where to make changes.
Overview
The library is built around a single class template, nlohmann::basic_json. A
basic_json value is a node in a tree of JSON values. All other components either create such a tree from an input
(parsing), write a tree to an output (serialization), or give access to it (iterators, JSON Pointer, conversions).
flowchart LR
input[/"input<br>(string, stream,<br>iterator range, file)"/]
ia["input adapter"]
lexer["lexer"]
parser["parser"]
breader["binary_reader"]
sax["SAX interface"]
value[("basic_json<br>value tree")]
serializer["serializer"]
bwriter["binary_writer"]
oa["output adapter"]
output[/"output<br>(string, stream,<br>vector)"/]
input --> ia
ia --> lexer --> parser --> sax
ia --> breader --> sax
sax --> value
value --> serializer --> oa
value --> bwriter --> oa
oa --> output
- JSON text is read by an input adapter, tokenized by the lexer, and turned into SAX events by the parser.
- Binary formats (BJData, BSON, CBOR, MessagePack, UBJSON) are read by an input adapter and turned into the same SAX
events by the
binary_reader. - A SAX consumer receives the events. The one used by
parsebuilds abasic_jsonvalue tree. - The
serializer(JSON text) or thebinary_writer(binary formats) writes a value tree to an output adapter.
Source layout
The public headers are in include/nlohmann:
json.hppdefines classbasic_json.json_fwd.hppcontains forward declarations.adl_serializer.hpp,byte_container_with_subtype.hpp, andordered_map.hppdefineadl_serializer,byte_container_with_subtype, andordered_map.json_view.hppis a separate, optional header that definesbasic_json_documentandbasic_json_view, a flat-index, read-only, non-owning way to look at a parsed JSON text; see Zero-copy JSON views. It builds onjson.hppinternals (it requires the same library version) and has its owndetail/view/subdirectory.
Everything else lives in detail/ and namespace nlohmann::detail, which is not part of the public API. Paths
below are relative to include/nlohmann.
| Component | Location |
|---|---|
| Value type enumeration | detail/value_t.hpp |
| Input adapters | detail/input/input_adapters.hpp |
| Lexer | detail/input/lexer.hpp, detail/input/number_parse.hpp, detail/input/string_scan.hpp |
| Parser | detail/input/parser.hpp |
| SAX interface and DOM builders | detail/input/json_sax.hpp |
| Binary format readers | detail/input/binary_reader.hpp |
| JSON serializer | detail/output/serializer.hpp, detail/conversions/to_chars.hpp |
| Binary format writers | detail/output/binary_writer.hpp |
| Output adapters | detail/output/output_adapters.hpp |
| Iterators | detail/iterators/ |
| Conversions from/to arbitrary types | detail/conversions/from_json.hpp, detail/conversions/to_json.hpp |
| JSON Pointer | detail/json_pointer.hpp |
| Exceptions | detail/exceptions.hpp |
| Type traits and C++ feature backports | detail/meta/ |
| Macros | detail/macro_scope.hpp, detail/macro_unscope.hpp, detail/abi_macros.hpp |
The single-header version single_include/nlohmann/json.hpp
is generated from these files with make amalgamate and must not be edited by hand. The same command also generates
single_include/nlohmann/json_view.hpp
from json_view.hpp and detail/view/.
Template parameters
basic_json is parameterized by the types it uses to store values and to convert from and to other types:
| Template parameter | Default | Used for |
|---|---|---|
ObjectType |
std::map |
objects, see object_t |
ArrayType |
std::vector |
arrays, see array_t |
StringType |
std::string |
strings and object keys, see string_t |
BooleanType |
bool |
Booleans, see boolean_t |
NumberIntegerType |
std::int64_t |
signed integers, see number_integer_t |
NumberUnsignedType |
std::uint64_t |
unsigned integers, see number_unsigned_t |
NumberFloatType |
double |
floating-point numbers, see number_float_t |
AllocatorType |
std::allocator |
allocating objects, arrays, strings, and binary values |
JSONSerializer |
adl_serializer |
conversions from/to other types, see adl_serializer |
BinaryType |
std::vector<std::uint8_t> |
binary values, see binary_t |
CustomBaseClass |
void |
an optional base class, see json_base_class_t |
The library provides two specializations:
jsonuses all default template arguments.ordered_jsonusesordered_mapasObjectTypeto keep the insertion order of object keys.
The requirements on the template arguments are listed in Template Parameter Requirements.
Value storage
Each basic_json value stores its content as a tagged union: an enumeration value_t
names the type of the value, and a union json_value holds the value itself. Both are members of the nested struct
data, which is the only data member m_data of basic_json:
struct data
{
/// the type of the current element
value_t m_type = value_t::null;
/// the value of the current element
json_value m_value = {};
};
data m_data = {};
with
enum class value_t : std::uint8_t
{
null, ///< null value
object, ///< object (unordered set of name/value pairs)
array, ///< array (ordered collection of values)
string, ///< string value
boolean, ///< boolean value
number_integer, ///< number value (signed integer)
number_unsigned, ///< number value (unsigned integer)
number_float, ///< number value (floating-point)
binary, ///< binary array (ordered collection of bytes)
discarded ///< discarded by the parser callback function
};
union json_value {
/// object (stored with pointer to save storage)
object_t *object;
/// array (stored with pointer to save storage)
array_t *array;
/// string (stored with pointer to save storage)
string_t *string;
/// binary (stored with pointer to save storage)
binary_t *binary;
/// boolean
boolean_t boolean;
/// number (integer)
number_integer_t number_integer;
/// number (unsigned integer)
number_unsigned_t number_unsigned;
/// number (floating-point)
number_float_t number_float;
};
Objects, arrays, strings, and binary values are allocated on the heap with AllocatorType, and the union only stores a
pointer to them. This keeps a basic_json value small: one pointer-sized union and one byte for the type. The class
maintains the invariant that the pointer matching m_type is never null; assert_invariant() checks it with
runtime assertions.
Node index of JSON views
A basic_json_document (see Zero-copy JSON views)
does not build a tree of values. Its parser
(detail/view/builder.hpp)
writes a flat array of 16-byte nodes, one per value and one per object key, in document order. A
basic_json_view is a pointer to the document and a pointer to one node. The layout
is struct node in
detail/view/node.hpp
(the numbers are bit offsets, 32 bits per row):
packet-beta
0-7: "kind"
8-15: "flags"
16-31: "extra"
32-63: "off"
64-95: "len"
96-127: "next"
| Bytes | Field | Type | Contents |
|---|---|---|---|
| 0 | kind |
uint8_t |
the type, numbered as value_t: 0 null, 1 object, 2 array, 3 string, 4 boolean, 5 signed integer, 6 unsigned integer, 7 float |
| 1 | flags |
uint8_t |
bits 0-1: where a string's bytes are (0: the source text, 1: the buffer of decoded strings, for strings with escapes); bit 2: the value of a boolean |
| 2-3 | extra |
uint16_t |
numbers: the number of integer digits (low byte) and fraction digits (high byte), 255 for more; objects: the number of their hash index (1-based), or 0; otherwise 0 |
| 4-7 | off |
uint32_t |
where the value starts: the first byte after a string's opening quote (or its position in the buffer of decoded strings), the first byte of a number or literal, the bracket of an array or object |
| 8-11 | len |
uint32_t |
strings: the length after decoding; floats and literals: the length of the token; arrays and objects: the number of elements |
| 12-15 | next |
uint32_t |
arrays and objects: the number of nodes of the subtree, including the node itself |
- Integers keep their converted value in bytes 8-15 instead of
lenandnext; the length of their token follows from the number of digits inextra(and the sign). Non-negative integers are unsigned integers, as withparse. - Floats keep only their token. The digit layout in
extralets the conversion read the digits without scanning the token again, and only when the value is read. - Object members are the node of the key (a string) followed by the nodes of the value.
- Navigation needs no pointers: the elements of an array or object follow its node, and the node after a value's
subtree is
nextnodes further for an array or object, and the next node otherwise (document_data::after). Views step from element to element this way and skip whole subtrees in constant time. - Offsets are 32 bits wide, so a document is limited to 4 GiB (
out_of_range.416). - Large objects (128 members or more) get a hash index after parsing
(
detail/view/object_index.hpp): an open-addressing table whose slots hold the distance from the object's node to a key's node, so that a lookup does not compare every key. The object'sextraholds the number of its table. Only 65,535 tables fit intoextra; objects beyond them are searched linearly.
For example, #!json {"a": [1, 2.5]} becomes five nodes. Each node's elements follow it, and next leads from an
array or object past its subtree:
flowchart LR
n0["0: object<br>len 1, next 5"]
n1["1: key a"]
n2["2: array<br>len 2, next 3"]
n3["3: unsigned integer 1"]
n4["4: float 2.5"]
e(["end"])
n0 --> n1 --> n2 --> n3 --> n4 --> e
n0 -. next .-> e
n2 -. next .-> e
| Node | kind |
extra |
off |
len |
next |
|---|---|---|---|---|---|
| 0 | 1 (object) | 0 | 0 | 1 | 5 |
| 1 | 3 (string) | 0 | 2 | 1 | 0 |
| 2 | 2 (array) | 0 | 6 | 2 | 3 |
| 3 | 6 (unsigned integer) | 0x0001 | 7 | - | - |
| 4 | 7 (float) | 0x0101 | 10 | 3 | 0 |
All flags are 0. The integer's bytes 8-15 hold its value, 1; its extra says it has one digit. The float's extra
says it has one integer and one fraction digit, and its len is that of the token 2.5.
Input adapters
Input is read via input adapters that abstract a source. Every input adapter provides this interface:
/// the type of the characters in the input
using char_type = ...;
/// read a single character; returns std::char_traits<char_type>::eof() at the end of the input
typename std::char_traits<char_type>::int_type get_character();
/// read up to count * sizeof(T) bytes into dest and return the number of bytes read
/// (used by the binary readers)
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1);
The lexer detects two optional extensions at compile time. Only iterator_input_adapter provides them, and only for
random-access input of single-byte characters:
supports_seek,get_consumed_count(), andcopy_consumed_range()let the lexer reconstruct already consumed input for error messages instead of copying every character it reads.supports_bulk_scan,bulk_data(),bulk_remaining(), andbulk_skip()let the lexer scan strings directly in contiguous memory, several bytes at a time.
The function input_adapter picks the right adapter for the argument passed to parse, accept, sax_parse, or the
from_* functions:
iterator_input_adapterreads from an iterator range, which also covers strings, containers, and pointers.wide_string_input_adapterreads from ranges ofwchar_t,char16_t, orchar32_tand converts them to UTF-8. It cannot be used for binary formats; itsget_elements()throws.input_stream_adapterreads from astd::istream.file_input_adapterreads from astd::FILE*.
SAX interface
The parser does not build values itself. It reports what it reads as events to a SAX
consumer, which implements the interface json_sax: null, boolean, number_integer,
number_unsigned, number_float, string, binary, start_object, key, end_object, start_array, end_array,
and parse_error.
The library comes with two consumers in detail/input/json_sax.hpp:
json_sax_dom_parserbuilds abasic_jsonvalue tree.parseuses it.json_sax_dom_callback_parserdoes the same, but calls a parser callback for each event, which can skip values.parseuses it when a callback is given.
The binary_reader emits the same events for binary formats, so sax_parse works
with a user-defined consumer for JSON and for all binary formats alike.
Output adapters
Output is written via output adapters:
void write_character(CharType c);
void write_characters(const CharType* s, std::size_t length);
The serializer (used by dump and operator<<) and the
binary_writer (used by the to_* functions) write to one of these adapters:
output_vector_adapterappends to astd::vector.output_stream_adapterwrites to astd::ostream.output_string_adapterappends to a string.
Value conversion
Values are converted from and to other types with the JSONSerializer template parameter. The default,
adl_serializer, calls the free functions
template<class T>
void to_json(basic_json& j, const T& t);
template<class T>
void from_json(const basic_json& j, T& t);
found by argument-dependent lookup. The library defines them for standard types in detail/conversions; users add them
for their own types, see Arbitrary Type Conversions. The
serialization macros generate these functions.
Additional features
- JSON Pointer (class
json_pointer) addresses values inside a tree. It is also the basis of JSON Patch. - Binary formats are read by
binary_readerand written bybinary_writer. - A custom base class can add members to every
basic_jsonvalue. - Serialization macros generate
to_jsonandfrom_jsonfunctions for user-defined types.
Details namespace
Namespace nlohmann::detail contains all implementation details. It is not part of the public API and may change in any
release. Besides the components above, it contains:
- type traits to detect the capabilities of user-defined types (
detail/meta/type_traits.hpp), - backports of C++14/17 features to C++11 (
detail/meta/cpp_future.hpp), and - helpers such as
string_concatandstring_escape.