fopen() of the CSV output was not checked, so an unwritable directory made
fprintf() write to a null FILE*; the round count is parsed with strtol and
clamped instead of atoi.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Hash the archives in chunks instead of reading them into memory, download
with a timeout, extract into a temporary directory that is renamed into
place only after success (a half-extracted directory was trusted forever),
and split CXX and CC into arguments so that values like 'ccache g++' work.
Note at the pins that a SHA-256 mismatch of a GitHub tag archive means
that GitHub regenerated it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
tests/benchmarks/json_view/ holds a comparison with other
libraries, answering the question users ask when they pick one.
It is not built by CMake and not run by CI.
- bench_view.cpp: parse, traverse, select and dump of twitter,
citm_catalog, canada, jeopardy, a tweet and an RPC request,
against json_view, yyjson, simdjson, Boost.JSON and json::parse,
all agreeing and run interleaved.
- bench_corpus.cpp: parse, traverse and dump of any file list.
- bench_edit.cpp: parses, edits through each library's own API,
and serializes; all outputs must describe the same value.
- compare.py: builds both programs with system or pinned,
SHA-256-checked downloads, and writes results with the commit,
CPU, OS, compiler and library versions.
- README.md and a workflow that runs on demand or via a PR label.
Fairness fixes folded in: each engine runs once untimed before
each timed round, so the next one no longer pays for the
previous one's cleanup; dump is also compared with source numbers
against yyjson's raw-number mode; a "simdjson DOM (fresh)" column
and a reused-document column for json_view were added;
compare.py's paths are absolute; downloads use a .part file and a
failed checksum removes the archive.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The dump and comparison tests moved to their own file in the dump pull
request; the tests of the fast dump() join them there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The MinGW linker of the Windows clang jobs cannot link object files with more
than 32767 sections ("relocation truncated to fit: IMAGE_REL_AMD64_REL32
against `.rdata'"). unit-json_view_edit.cpp has 35409 sections since the fast
dump instantiates more code for editable documents, so its test cases "views
and values", "deeply nested values", "pointers below a null value", and
"strings of other documents are checked" move into
unit-json_view_edit_values.cpp. The helper exception_of_call() that both
files use moves into json_view_test_helpers.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A float node of an image loaded with image_check::bounds can hold a
token whose bytes are not digits, so its value need not have the digit
count recorded in the node. write_double_at() passed that count to
write_short_decimal(), which indexes its table of powers of ten by it:
an assertion failure in debug builds, an out-of-bounds read in release
builds. Use the counted overload only if the value has exactly that many
digits, otherwise count them.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The default dump() (no indentation, no ensure_ascii) gets its own
writer that makes the same walk and produces the same output:
- the write position stays in a local variable instead of a
member, so the compiler keeps it in a register across stores
through aliasing char pointers;
- strings and number tokens are copied with fixed-size 32-byte
moves wherever enough source bytes remain, instead of one
memcpy call per token;
- the innermost open container lives in local variables; a stack
that starts as a local array of 32 entries holds the rest;
- unedited documents are walked through the node array in order,
and integer tokens are read from the source directly.
On top of that, float tokens of at most 15 significant digits are
written straight from their digits via zmij::to_shortest() and
write_shortest(), without converting to a double and back: such
decimals are farther apart than a double's rounding interval, so
the token's digits are the double's shortest digits. Tokens of
16+ digits, or edited values, still go through decimal_to_float().
The view's own NEON write_decimal() is removed in favor of the
shared writer, and the dump output now grows in 64 KiB steps
instead of being resized to its estimate at once.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
load_result() exists only with exceptions; with JSON_NOEXCEPTION the test
loads the image with the check directly (a failed check aborts).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
bugprone-suspicious-stringview-data-usage flags data() of the source view
under C++17, although the length is passed and checked by the REQUIRE
before it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The full check scanned the contents of every string node and every float token over its own range, so many nodes pointing to one large range made load() quadratic. The structural walk now only collects the ranges; after it, each kind is sorted and checked once per distinct range. Ranges of one kind that overlap without being identical are rejected, as save() never writes them (nodes that share a value share the whole range). The cost is linear in the size of the image plus sorting the ranges.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Move the raw string literal out of the CHECK macro (MSVC preprocessor),
cast 64-bit header fields to std::size_t (C4244 on 32-bit), and give the
image_check section in load.md a real anchor for the mkdocs strict build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().
load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.
image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.
Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.
New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.
A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups find the first member of a repeated key again, so set(object, key,
value) assigns that member (keeping the key where it is) and drops the
others, as documented before, and a view taken from object[key] shows the
new value. This undoes the code and documentation changes of 5cbb8e6d6.
Lookups in edited objects (navigation of editable documents) stop at the
first match, like those of parsed objects.
The tests that commit added expect the first member from lookups now. In the
seeded differential test, the documents with repeated keys repeat them after
the real members; the reference holds the first members, which the edits
address and the lookups are compared with, while materialize() is compared
with what parse() makes of the document's dump.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An editable document reads the last member of a repeated key (as the
read-only document does), but set(object, key, value) assigned the first
one: a view taken from object[key] before the call did not show the new
value. It now assigns the last member's value, keeps the key at the
position of its first occurrence (where materialize() puts it), and drops
the other members.
The seeded differential test now also edits documents whose objects
repeat keys and compares every lookup (operator[], at, find, value,
contains, count, and JSON pointers) with the parsed basic_json value.
A targeted test covers duplicates before and after edits that move the
object, in large objects with an index, in moved arrays, and in values
copied from other documents.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Replacing an array or object that spans several nodes by a scalar looks
up its parent from the root (find_parent), which is linear in the size
of the document: setting every element of a 20,000 element array took
more than half a second. The parent only has to switch to links once;
afterward the extent of the replaced value no longer matters. Nodes that
an entry of a moved sequence links to are now marked (node_flags::linked),
and assign() skips the lookup for them.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
set(json_pointer, value) turned a null parent into an object for every
last reference token, so "/a/0" and "/a/-" below {"a":null} made
{"a":{"0":1}}, where basic_json's operator[](json_pointer) makes an
array. A null parent now becomes an array for "-" and for digit tokens
(filled with nulls up to the index) and an object otherwise. An invalid
index is reported before the parent changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
append_text() doubled the capacity of the arena and rejected the result
if it exceeded 4 GiB, so that an arena of more than 2 GiB could not grow
even though the 32-bit offsets of nodes address 4 GiB - 1 bytes. The
capacity is now clamped to that limit (text_capacity()), and an append
is rejected only if the bytes themselves do not fit.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An editable document only holds valid UTF-8, but copy_scalar() copied the
strings and keys of a view of another document unchecked. A document of
a weaker check (or a borrowed text that changed after parsing) could
therefore bring ill-formed UTF-8 into it. The copy is checked now, with
the error that dump() reports for the string; copies within the same
document stay unchecked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Counting and copying the nodes of a view or a basic_json value into an
editable document recursed once per nesting level, so that a deeply
nested value overflowed the stack. The four functions now walk the value
with an explicit stack, as materialize() does.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.
Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.
New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.
Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).
A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The dump and comparison tests moved to their own file in the dump pull
request; the hash index tests of this branch join them there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.
Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.
On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.
Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.
simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The MinGW linker of the Windows clang jobs cannot link object files with more
than 32767 sections ("relocation truncated to fit: IMAGE_REL_AMD64_REL32
against `.rdata'"). unit-json_view.cpp reaches that limit as the stack
grows, so its "json_view dump" and "json_view comparison" test cases move
into unit-json_view_dump.cpp. The test generator and has_duplicate_keys()
that both files use move into json_view_test_helpers.hpp.
The new file mentions JSON_HAS_CPP_17, so it is built for C++17 like the file
it was split from.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
dump() writes every member and == resolves duplicate keys as parse()
does, whereas lookups find the first member of a duplicate key.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The source extent of a value was read from the next node, falling back to the rest of the document when that node held a decoded string. dump() of a small value could thus allocate a buffer as large as the document. Skip a few such nodes, cap the fallback estimate, and shrink a buffer that is much larger than its output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>