With the nested struct object_index, clang 4 (and, by the same bug, the
clang 3.x of ci_test_compilers_clang) rejects the explicitly noexcept
defaulted constructor: "default member initializer for 'indexes' needed
within definition of enclosing class 'document_data' outside of member
functions". Nothing depends on the constructor being noexcept, so let it
take the implicit exception specification.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast on Linux x86-64 rejects
static_cast<std::size_t>(key_hash(...)): the call returns a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Store the hash in a variable
and cast that, which GCC does not report.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast on Linux x86-64 rejected
static_cast<std::uint64_t>((std::numeric_limits<std::size_t>::max)()),
as both are the same type there. Compare without the cast: std::size_t
converts to std::uint64_t implicitly on every platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
constant) for `TrailingCommas && cur() == ']'` and the like when the
option is off. Route the template arguments through a static enabled()
function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
single_include/, which does not contain the internal
nlohmann/detail/view headers. Build that test only with the multiple
headers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The default dump() (no indentation, no ensure_ascii) gets its own
writer: the same walk and output, with the write position in a local
variable (stores through char pointers would otherwise force a reload
of the buffer's members after each one), strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left (the buffer keeps 64 bytes of slack), and decoded
strings copied in runs up to the next quote, backslash, or control
character. Documents that are not edited are walked through the node
array in order, so that a frame only needs the end of its container,
and integer tokens are read from the source directly. The innermost
open container is kept in local variables, and the stack holds only
the ones around it; the stack starts in a local array of 32 and moves
to the heap only for deeper nesting (its address does not escape, so
its pointers stay in registers). Dumps of shallow documents thus
allocate only the output, whose first size includes the slack, so it
does not grow just before the end.
The long copies are out of line: otherwise, the compiler merges the
fixed-size moves into the same library call.
These techniques come from the prototype; the writer lost them when the
view was split into pull requests, which made dump() 2 to 3 times
slower.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).
- save() writes an edited document in its current state, in document
order (floats that are not finite become null, as in dump()); the
same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
load(vector&&) keeps it without a copy. The nodes are copied (aligned,
and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
bounds, UTF-8, strings of the source, number tokens and their values);
bounds checks structure and bounds, so that reading and serializing
stay safe; none trusts the image.
A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.
As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.
The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.
Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names
The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.
Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Test edits of an empty document and assignments through a view of a
value that is no longer part of the document; copy the entries of a
block through deref() (one path for links and values); mark the
4 GiB limit and the returns of find_parent() that no document reaches.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:
- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)
Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.
Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.
Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:
- node flags `edited` (a string or number token in the edit arena),
`moved` (the elements of an array/object live in a separate sequence),
and `is_new` (no source position), and link nodes (kind_link) that
stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
values, and the edit arena
materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.
Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.
Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.
Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.
json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.
Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For doubles with at most 19 digits, the
value is now read from that layout: the digits eight at a time, without
scanning the token, and rounded with Clinger's fast path where both
operands are exact, else with the Eisel-Lemire algorithm (which needs no
fallback for up to 19 digits). Both round correctly, so the values are
those of parse(); other tokens and types keep the library's conversion.
get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.53 -> 0.86 GB/s.
Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22) to the
bit-for-bit comparison with parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.
Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.
The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.
Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get<T>() of arithmetic types is inlined down to the conversion, so that
its checks of the node kind merge with those of the caller, and reading
an integer needs no call. Traversing every value: citm_catalog -6%,
marine_ik -5%, numbers and twitter -3%, mesh -2.5%, canada -1% (and
more above the float conversion from the digit layout: citm_catalog
-14%, marine_ik -11%).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_string() and number_token() return braced lists; the test compares
floats by their bit patterns instead of with memcmp, uses std::any_of, and
marks a fixed seed, a default member initializer (needed by GCC's
-Weffc++), and a string search.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains get<T>(), get_to(), value() with keys and JSON
pointers, and operator[], at(), and contains() with JSON pointers, plus
two functions basic_json has no counterpart for:
- get_string(): the string without a copy (a string_view into the source,
or into the decoded strings for strings with escapes)
- number_token(): the text of a number as it appears in the source
get<T>() converts arithmetic types, strings (also string_view_t),
std::nullptr_t, std::vector, maps with string keys, and views directly;
floats are converted from the digit layout recorded by the parser with the
library's conversion chain, so the values are bit-identical to parse().
Other types, including user types with from_json(), go through
materialize().
The exceptions are those of basic_json, message included. Where const
basic_json has undefined behavior (a missing key or an index out of range
with operator[] and a JSON pointer), the result is a discarded view;
value() returns the default wherever basic_json catches out_of_range, and
contains() never throws. Array indices of JSON pointers follow
json_pointer's rules (parse_error.106/109, out_of_range.404/410).
Tests compare the conversions of 2,000 generated documents, 20,000 float
tokens (double and float, bit for bit), and every JSON pointer of 1,000
documents with basic_json, and the exceptions for malformed pointers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Marks the default initializer of the item's index string (needed by GCC's
-Weffc++) and, in the test, an escaped literal and a comparison of find()
with end(), which is what the test is about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
detail::json_pointer_access returns the reference tokens of a pointer, so
that code resolving pointers without a basic_json value (such as the
zero-copy view) does not have to parse to_string() again. No change in
behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>