- API pages for operator[], at, front, back, find, contains, count,
begin, end, cbegin, cend, items, and type_name of basic_json_view,
linked both ways with the basic_json pages
- the feature page and size() describe document order and duplicate
keys
- the examples show when the view helps: reading a few fields of a large
text, probing optional members, and members in source order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains the read-only access functions of basic_json:
operator[] and at() with keys and indices, front(), back(), find(),
contains(), count(), begin()/end(), items() (with structured bindings from
C++17 on), and type_name(). They throw the exceptions (ids and messages)
that the const functions of basic_json throw; where basic_json has
undefined behavior (operator[] with a missing key or an index out of
range, front()/back() of an empty container), the view returns a
discarded view or throws invalid_iterator.214.
Objects are iterated in document order, and all members are visited. With
duplicate keys, lookups find the first member, so that a lookup can stop
at the first match; parse() keeps the last value. Keys of up to 16 bytes
are compared with two overlapping loads instead of memcmp, and most keys
are rejected by their length alone, from the index.
The iterators and items live in detail/view/iterator.hpp, the lookups in
detail/view/lookup.hpp. Tests compare every element and member of 2,000
generated documents with ordered_json, keys of every length around the
load sizes, the exceptions against const basic_json, and the iterators.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json::type_name() now calls detail::value_type_name(value_t), so
that code which reports types without a basic_json value at hand, such as
the zero-copy view, uses the same names in its exception messages. No
change in behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "check" job (Check amalgamation) runs develop's amalgamate.py and read
all configurations from the develop checkout, where config_json_view.json
does not exist until this stack lands, so it failed with
FileNotFoundError. Read that configuration from the pull request's
checkout; the tool itself stays develop's.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: the helpers that compare the exceptions of
json_document::parse() and json::parse() catch them outside a
CHECK_THROWS, so with JSON_NOEXCEPTION the first parse error aborted
the test. Compile those comparisons only with exceptions, as
unit-class_parser.cpp does.
- ci_test_gcc: -Werror=unused-result for CHECK_THROWS_AS(json_document::
parse(...)); assign the result to a dummy document.
- ci_test_compilers_clang (3.6): `const json_view invalid;` needs a
user-provided default constructor there (CWG 253); value-initialize it.
- ci_test_single_header: json_view.hpp now exists as a single header and
contains the internal view headers, so unit-json_view_builder.cpp
includes it instead of the detail headers in that mode, and the test
is built again with the single header.
- Regenerate single_include/nlohmann/json_view.hpp for the builder change
merged from json-view/08-view-builder.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A new section describes the 16-byte node: its fields, how integers,
floats, and object members are stored, how views navigate without
pointers, and a worked example. The feature page and the pages of
basic_json_document and node_count link to it where they mention the
16 bytes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- the input dispatch takes byte ranges by const reference and reads the
size once (which also settles a finding of the static analyzer); input
adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
benchmarks_view.cpp adds ViewParse, ViewRead (a reused document),
ViewParseIndented, ViewAccept, and ViewMaterialize on the files of
ParseString, so that each row can be read against the json::parse row of
the same file; benchmarks.cpp gains Accept (json::accept) as the
counterpart of ViewAccept. The view benchmarks are built only if the
header directory has json_view.hpp, so that older versions can still be
benchmarked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for basic_json_document and basic_json_view, one per member,
and for the four aliases, each with an example
- features/json_view.md: the problem the view solves, ownership and
lifetime, what matches basic_json::parse() and what differs, and when
to choose json, ordered_json, SAX, or the view
- the examples show why one would use the view, not only how: borrowed
vs. owned input, reading a few fields and materializing one subtree,
reusing a document across many messages
- registered in the mkdocs navigation, llms.txt, the docset, the
exceptions page (out_of_range.416), architecture.md, the integration
page, and the README; the yyjson credit is added to the README and
license.md
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
fuzzer-parse_json_view.cpp checks for every input that
json_document::accept agrees with json::accept, that an accepted input
materializes to the value json::parse returns, and that a rejected input
makes both throw the same exception with the same message. It is built
like the other fuzzers (tests/Makefile, and the root Makefile's
fuzz_testing_json_view target, which starts from the JSON test corpus) and
listed in tests/fuzzing.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The nlohmann.json module includes json_view.hpp in its global module
fragment and exports basic_json_document, basic_json_view, and the four
aliases next to basic_json, json, and ordered_json. features/modules.md
lists them, and tests/module_cpp20 parses a document, so that a missing
export fails the ci_module_cpp20 job.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The new single header goes through the same checks and install steps as
json.hpp and json_fwd.hpp:
- cmake/ci.cmake: ci_test_amalgamation regenerates, formats, and compares
json_view.hpp as well
- check_amalgamation.yml: the pull request check does the same; it runs
develop's tools, so it needs config_json_view.json on develop first
- meson.build: installs single_include/nlohmann/json_view.hpp
- gen_bazel_build_file.cmake, BUILD.bazel: json_view.hpp joins the
single-header target; the glob of the other target already covers the
new headers
- labeler.yml: an "aspect: json_view" label for the header, its
detail/view headers, tests, and documentation
The CMake install rules for include/ and single_include/, the REUSE
catch-all, and Package.swift cover the new files without changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
unit-json_view.cpp: type queries, size, and empty against basic_json;
materialize() against parse() (json and ordered_json, generated documents,
duplicate keys, 100,000 levels of nesting, parent pointers with
JSON_DIAGNOSTICS); parse errors and their messages equal to parse() for
malformed inputs and all option combinations; the overflow of a float
document (1e39, 3.4028236e38) as in parse(); NUL and BOM; borrowed and
owned inputs (strings, C strings, literals, vectors, string_view, streams,
wide strings, parse_copy, and iterator ranges over pointers, vectors,
strings, and lists); reuse with read(); moves; shrink_to_fit() of the index
and of the decoded strings; source offsets.
unit-json_view_macros.cpp includes the header without
JSON_TEST_KEEP_MACROS, as users do: the view must not depend on the macros
json.hpp undefines, and must not leak its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:
- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
inputs, owning rvalue strings, streams, and other inputs), parse_copy,
accept, read, root, is_discarded, source, owns_source, node_count,
memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
materialize (the value parse() would produce, built by the same SAX
handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
ordered_json_view
A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.
The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
constant) for `TrailingCommas && cur() == ']'` and the like when the
option is off. Route the template arguments through a static enabled()
function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
single_include/, which does not contain the internal
nlohmann/detail/view headers. Build that test only with the multiple
headers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Numbers of more than 19 digits at the boundary of the largest double are
decided by the exact comparison with the midpoint in the overflow check.
The check uses the floating-point type of the document: with float, the
view rejects what parse() rejects (1e39, 3.4028236e38, the midpoint between
the largest float and 2^128), and double documents are not affected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- tables as std::array; the frames of the first 64 levels stay a C array
(not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
second table
- the parse failure is private, with an accessor; the special member
functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.
Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.
Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:
- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
its own prefix (json.hpp undefines its own at its end); the throw macro
honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
(checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
(after yyjson) and the library's SWAR and UTF-8 checks, independent of the
byte order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.
json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.
Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast rejected static_cast<std::size_t>(next() % n):
on 64-bit Linux std::uint64_t and std::size_t are the same type, while
the cast is needed where std::size_t is 32 bits wide. Draw the sizes from
a 32-bit value instead, which converts to std::size_t implicitly on every
platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.
The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The strtold fallback, which is left only for long double formats that
are not binary64 (x87, binary128), substituted the first byte of the
locale's decimal point for '.'. Under a locale whose decimal point is
longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B),
strtold stopped there and the value was truncated at the decimal point.
A longer decimal point is now put into a copy of the token.
The test "locale with a multi-byte decimal point" now compares the long
double values with those of the "C" locale; with x87 long doubles it
failed before.
Fixes#5660.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:
- The token is split into sign, significand w (at most 19 digits), and
decimal exponent q, using the positions of the decimal point and the
exponent that the scanners already recorded, so no character is
classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
an exact big-integer comparison with the midpoint between the two
candidates (the digit comparison of fast_float, simplified).
This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.
convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.
Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.
Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().
Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.
Fixes#5652.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.
Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.
Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).
Fixes#5666.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add JSON_NO_UDLS to leave out the user-defined string literals
The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mention JSON_NO_UDLS in the list of exported module symbols
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the user-defined string literals to <nlohmann/json_literals.hpp>
Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.
The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test
- Suppress clang-tidy misc-header-include-cycle on the intentional mutual
include of json.hpp and json_literals.hpp.
- Use operator"" _json with a space for GCC 4.8 in the test's detection
aliases, as the header does.
- Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Declare the literal operators through a local macro
The GCC 4.8 spacing condition was repeated for both operator definitions
and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix)
now selects operator""##suffix or operator"" suffix in one place and is
undefined at the end of json_literals.hpp.
Suggested by gregmarr in review.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When copying a value nested deeper than 128 levels, and an allocation
fails while an inner array or object is being copied, the partially
built copy ended up with an element typed array/object but holding a
null pointer. That element was already a fully constructed member of
its parent's container, so destroying the parent during stack
unwinding dereferenced the null pointer (release builds) or failed
assert_invariant() (debug builds), instead of letting std::bad_alloc
reach the caller.
copy_iteratively() set a pending worklist element's type right after
popping it, before the next loop iteration created its container in
copy_array_level()/copy_object_level(). Move that type assignment into
those two functions, right after the container is successfully
created, and drop the premature one in copy_iteratively(), so a
half-built element stays a null value - as copy_shallow()'s comment
already promised - until it can safely hold one.
Add a regression test to tests/src/unit-allocator.cpp that copies a
value nested 130 levels deep (both arrays and objects, with a
std::map- and an ordered_map-backed object_t) and fails every
allocation of the copy in turn: each attempt must throw std::bad_alloc
without crashing, and the source must stay unchanged.
Fixes#5640.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
For an object type whose comparator treats unequal keys as equivalent
(for example a std::map with a case-insensitive comparator),
compare_iteratively() looked up a mismatched left key in the right
object with find(), which uses the object's own comparator, and
accepted whatever entry it found without checking that the keys are
actually equal. A case-insensitive comparator then found "KEY" for
"key", so two objects nested past the recursion bound (or at every
depth with JSON_NO_THREAD_LOCAL) could compare equal even though the
object type's own operator== - and basic_json itself, below the bound
- consider them different.
Accept the found entry only if its key equals (not just compares
equivalent to) the looked-up key.
Fixes#5655.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json(first, last) treated value_t::binary like the structured
types (array, object) in the range check, so it always copied the
whole binary value regardless of the iterators, even for an empty
range such as (b.end(), b.end()). The other primitive types (number,
boolean, string) already reject such a range with
invalid_iterator.204, and erase(first, last) already does the same
for binary values, so this made the constructor inconsistent with
both. Move case value_t::binary into the group of checked primitive
types.
Also update the two matching passages in basic_json.md that describe
overload 7, and add a version-history note.
Fixes#5670.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.
The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.
Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.
Fixes#5643.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is
a null pointer constant, so it converts to a null const char*, and the
overloads taking const typename object_t::key_type& accepted it by
constructing a std::string from that null pointer, which is undefined
behavior (a crash with both libc++ and libstdc++). value(0, default_value)
had the same problem in C++11, where the object comparator is not
transparent.
Add deleted overloads for integral arguments to contains(), find()
(const and non-const), count(), and value() so that these calls are
compile errors in every supported language mode instead of crashing.
Calls with string, string_view, json_pointer, and size-typed element
access (at(), operator[](), erase()) are unaffected.
Fixes#5657.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.
calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.
Fixes#5675.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.
Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.
Fixes#5658.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Parsing from a std::istream crashed in two unusual but valid stream
states, both in input_stream_adapter:
- With eofbit in the stream's exceptions() mask, get_character() sets
eofbit via is->clear(), which throws std::ios_base::failure. While
that exception unwinds, ~input_stream_adapter() called clear() again
to reset eofbit, which is still set and still in the exception mask,
so it throws a second time out of the (implicitly noexcept)
destructor and std::terminate() is called. The destructor now only
calls clear() if a bit other than eofbit remains set, so the first
exception can propagate normally.
- For an std::istream without a stream buffer (rdbuf() == nullptr,
e.g. std::istream(nullptr)), the constructor stored the null
pointer without checking it, and get_character() dereferenced it.
input_adapter(std::istream&) now throws parse_error.101 for such a
stream, the same as it already does for a null FILE* or char*.
Added regression tests to unit-deserialization.cpp and, for the
JSON_PRECISE_STREAM_POSITION variant of get_character(), to
unit-precise-stream-position.cpp; both crashed before this fix.
Documented the two exceptions in parse.md and operator_gtgt.md.
Fixes#5646.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
from_json built its out_of_range.410 message with "..." + j.dump(). If the
unmatched value is (or contains) a string with invalid UTF-8, that dump()
itself throws type_error.316, so the caller got type_error.316 instead of
the documented out_of_range.410; such strings can reach get<Enum>()
unvalidated, e.g. from from_cbor()/from_msgpack(). With a custom string_t,
j.dump() returns that type, and "const char*" + string_t does not compile
unless the type happens to provide operator+, so the macro failed to
compile for such types.
Build the message with detail::concat(), which appends any type exposing
data()/size() and always yields a std::string, and dump with
error_handler_t::replace so building the message itself cannot throw.
Added regression tests: an invalid-UTF-8 case in the existing strict-enum
test in unit-conversions.cpp, and a strict-enum use with alt_string (the
custom string_t from unit-alt-string.cpp) to cover the compile failure.
Fixes#5667.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap
overload that forward to it) only exchanged m_data.m_type/m_data.m_value,
leaving each value's json_base_class_t subobject in place. This is
inconsistent with the copy and move constructors and copy assignment,
which all carry the base class along with the value, so after
a.swap(b) any metadata stored in a CustomBaseClass ended up attached to
the wrong value. Algorithms that mix swap() with moves, such as
std::sort, scrambled the metadata across the whole container.
Fix the member swap() to also exchange the json_base_class_t subobject
and extend the noexcept specifications of swap() and the friend swap()
accordingly.
Fixes#5653.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix to_msgpack() reading the inactive number union member
basic_json stores number_integer and number_unsigned in a union, and
number_unsigned_t only has to be at least as wide as number_integer_t
(with the default types, both are 64-bit and have the same
representation). When number_integer_t is narrower, write_msgpack()
read the wrong union member in two places:
- The number_unsigned case wrote number_integer's bits instead of
number_unsigned's, silently writing the wrong value whenever it
did not fit in number_integer_t.
- The number_integer case (non-negative branch) picked the encoded
width by comparing number_unsigned's bits, which is undefined
behavior, though the value written was still number_integer's, so
at worst a too-wide encoding was chosen.
Read the active member in both cases, like the other binary writers
(CBOR, UBJSON, BJData, BSON, BON8) already do.
Fixes#5644.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Cast number_integer to number_unsigned_t only once in to_msgpack()
Addresses review comment by @gregmarr.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Copy values before inserting an initializer list into an array
insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).
Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.
Fixes#5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the reserve_array helper in the initializer_list insert fix
The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.
Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take diff()'s fast path unless the object type reorders members
For every object type except an insertion-ordered one like ordered_map,
diff() no longer produced a member-by-member patch when target had a key
that sorts before a key the two objects share: it fell through to the
slow path, which removes every member of source and re-adds every member
of target, instead of just adding the new key.
#5465 added an order check to require the fast path to also reproduce
target's member order, needed because ordered_map's patch()-driven "add"
appends a new member at the end. The check compared the common keys'
order between source and target and also required that every added key
come after every common key in target's order ("new_keys_form_suffix").
The comment above it argued this check is always true for std::map, and
that reasoning is correct for the order of the common keys themselves,
but not for new_keys_form_suffix: a std::map iterates in sorted key
order, so a new key that sorts before an existing common key is
enumerated between common keys, making new_keys_form_suffix false even
though std::map's own key order does not need reordering at all - it
places every member itself, regardless of insertion history, so a
member-by-member diff already reproduces target's iteration order.
Only require the order check for an object type that keeps insertion
order, using the same detail::is_ordered_map trait the library already
uses to recognize such an object type in set_parent(). Every other
object type - std::map in key order, a hash map in an order its
operator== ignores - always takes the fast path.
Added a regression test to unit-json_patch.cpp: the issue's example now
yields a single "add" op for json, while ordered_json still takes the
slow path to reproduce target's member order.
Fixes#5639.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Only track the target key order in diff() for insertion-ordered objects
common_keys_target_order and new_keys_form_suffix are only read when object_t keeps its members in insertion order; skip building them otherwise. Addresses review comment by @gregmarr.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>