The "check" job (Check amalgamation) runs develop's amalgamate.py and read
all configurations from the develop checkout, where config_json_view.json
does not exist until this stack lands, so it failed with
FileNotFoundError. Read that configuration from the pull request's
checkout; the tool itself stays develop's.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: the helpers that compare the exceptions of
json_document::parse() and json::parse() catch them outside a
CHECK_THROWS, so with JSON_NOEXCEPTION the first parse error aborted
the test. Compile those comparisons only with exceptions, as
unit-class_parser.cpp does.
- ci_test_gcc: -Werror=unused-result for CHECK_THROWS_AS(json_document::
parse(...)); assign the result to a dummy document.
- ci_test_compilers_clang (3.6): `const json_view invalid;` needs a
user-provided default constructor there (CWG 253); value-initialize it.
- ci_test_single_header: json_view.hpp now exists as a single header and
contains the internal view headers, so unit-json_view_builder.cpp
includes it instead of the detail headers in that mode, and the test
is built again with the single header.
- Regenerate single_include/nlohmann/json_view.hpp for the builder change
merged from json-view/08-view-builder.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
constant) for `TrailingCommas && cur() == ']'` and the like when the
option is off. Route the template arguments through a static enabled()
function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
single_include/, which does not contain the internal
nlohmann/detail/view headers. Build that test only with the multiple
headers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast rejected static_cast<std::size_t>(next() % n):
on 64-bit Linux std::uint64_t and std::size_t are the same type, while
the cast is needed where std::size_t is 32 bits wide. Draw the sizes from
a 32-bit value instead, which converts to std::size_t implicitly on every
platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...))
because parse() is [[nodiscard]]; assign the result to a dummy json as the
other tests do. clang-tidy flagged longer.find('.') == npos with
abseil-string-find-str-contains; store the position in a variable first.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A new section describes the 16-byte node: its fields, how integers,
floats, and object members are stored, how views navigate without
pointers, and a worked example. The feature page and the pages of
basic_json_document and node_count link to it where they mention the
16 bytes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- the input dispatch takes byte ranges by const reference and reads the
size once (which also settles a finding of the static analyzer); input
adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
benchmarks_view.cpp adds ViewParse, ViewRead (a reused document),
ViewParseIndented, ViewAccept, and ViewMaterialize on the files of
ParseString, so that each row can be read against the json::parse row of
the same file; benchmarks.cpp gains Accept (json::accept) as the
counterpart of ViewAccept. The view benchmarks are built only if the
header directory has json_view.hpp, so that older versions can still be
benchmarked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for basic_json_document and basic_json_view, one per member,
and for the four aliases, each with an example
- features/json_view.md: the problem the view solves, ownership and
lifetime, what matches basic_json::parse() and what differs, and when
to choose json, ordered_json, SAX, or the view
- the examples show why one would use the view, not only how: borrowed
vs. owned input, reading a few fields and materializing one subtree,
reusing a document across many messages
- registered in the mkdocs navigation, llms.txt, the docset, the
exceptions page (out_of_range.416), architecture.md, the integration
page, and the README; the yyjson credit is added to the README and
license.md
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
fuzzer-parse_json_view.cpp checks for every input that
json_document::accept agrees with json::accept, that an accepted input
materializes to the value json::parse returns, and that a rejected input
makes both throw the same exception with the same message. It is built
like the other fuzzers (tests/Makefile, and the root Makefile's
fuzz_testing_json_view target, which starts from the JSON test corpus) and
listed in tests/fuzzing.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The nlohmann.json module includes json_view.hpp in its global module
fragment and exports basic_json_document, basic_json_view, and the four
aliases next to basic_json, json, and ordered_json. features/modules.md
lists them, and tests/module_cpp20 parses a document, so that a missing
export fails the ci_module_cpp20 job.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The new single header goes through the same checks and install steps as
json.hpp and json_fwd.hpp:
- cmake/ci.cmake: ci_test_amalgamation regenerates, formats, and compares
json_view.hpp as well
- check_amalgamation.yml: the pull request check does the same; it runs
develop's tools, so it needs config_json_view.json on develop first
- meson.build: installs single_include/nlohmann/json_view.hpp
- gen_bazel_build_file.cmake, BUILD.bazel: json_view.hpp joins the
single-header target; the glob of the other target already covers the
new headers
- labeler.yml: an "aspect: json_view" label for the header, its
detail/view headers, tests, and documentation
The CMake install rules for include/ and single_include/, the REUSE
catch-all, and Package.swift cover the new files without changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
unit-json_view.cpp: type queries, size, and empty against basic_json;
materialize() against parse() (json and ordered_json, generated documents,
duplicate keys, 100,000 levels of nesting, parent pointers with
JSON_DIAGNOSTICS); parse errors and their messages equal to parse() for
malformed inputs and all option combinations; NUL and BOM; borrowed and
owned inputs (strings, C strings, literals, vectors, string_view, streams,
wide strings, parse_copy, and iterator ranges over pointers, vectors,
strings, and lists); reuse with read(); moves; shrink_to_fit() of the index
and of the decoded strings; source offsets.
unit-json_view_macros.cpp includes the header without
JSON_TEST_KEEP_MACROS, as users do: the view must not depend on the macros
json.hpp undefines, and must not leak its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:
- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
inputs, owning rvalue strings, streams, and other inputs), parse_copy,
accept, read, root, is_discarded, source, owns_source, node_count,
memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
materialize (the value parse() would produce, built by the same SAX
handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
ordered_json_view
A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.
The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Numbers of more than 19 digits at the boundary of the largest double
are decided by the locale-aware fallback of the overflow check.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- tables as std::array; the frames of the first 64 levels stay a C array
(not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
second table
- the parse failure is private, with an accessor; the special member
functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.
Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.
Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:
- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
its own prefix (json.hpp undefines its own at its end); the throw macro
honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
(checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
(after yyjson) and the library's SWAR and UTF-8 checks, independent of the
byte order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.
json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.
Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.
The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).
The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.
json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
lexer::convert_number() converted float tokens with std::from_chars (when
available), Clinger's fast path, and the locale-aware strtod fallback, all
as lexer members. They are now free functions in number_parse.hpp:
- convert_float_fast(): std::from_chars, then Clinger's fast path, skipped
when the mantissa has too many significant digits
- convert_float_locale_aware(): strtof/strtod/strtold with the decimal point
of the current locale, retried when the locale changed (#5198)
so that other code converting JSON number tokens gets the same values. No
change in behavior; the lexer no longer includes <clocale> and <cstdlib>.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A header that builds on json.hpp (such as the planned json_view.hpp) must
not inline json.hpp: its single-header version would contain a second copy
of the library, and that copy would change with every library change.
The optional config key "external" lists include paths that are kept as
#include directives. Only the first directive per path is kept; repeated
ones are commented out, as the tool already does for inlined headers.
config_json_view.json uses it for json_view.hpp; the existing configs do
not set it, and json.hpp and json_fwd.hpp regenerate byte-identically.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:
- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
calls whose result was discarded. Check the result, like the other
resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
callbacks, which cannot throw but were not declared noexcept.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Look up the locale decimal point at conversion time, not lexer construction
The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.
token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.
Fixes#5198
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop the strtod retry loop when the decimal point is unchanged
convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.
Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Weffc++ errors in the #5198 locale test
GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the key type when rejecting non-string CBOR/MessagePack map keys
CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:
syntax error while parsing MessagePack object key: only string keys
are supported, but found nil; last byte: 0xC0
The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.
Refs #2766, #3381
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point the MessagePack key note to the spec's profile section
The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).
Write into a caller-owned vector instead, so nothing is reserved.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5559 added a 224-character line to cbor_tag_handler_t.md; the
documentation style check allows at most 160.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The line added in #5559 exceeded the 160-character limit enforced by
docs/mkdocs/scripts/check_structure.py, breaking the documentation build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests added by #5515 fail two ways on develop:
- clang-tidy reports the size() overrides of huge_string and huge_binary
(readability-convert-member-functions-to-static) and the non-const
test value (misc-const-correctness); mark them like the #5584 types
- clang with libstdc++ 10 cannot compile the file for C++17: the
std::filesystem::path conversion considered for huge_string, a class
derived from std::string, is ambiguous. Guard it with
JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same
reason, and define that macro before both test blocks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add BON8 support
Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.
The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.
The round-trip tests need the .bon8 files of json_test_data 3.2.0.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Address review comments
- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Select the BON8 float prefix by type
get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Rename a test variable that Flawfinder mistakes for read()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures
- compare the float in write_bon8_float with number_float_t constants,
so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BON8 strings in bulk from contiguous input
- copy the valid UTF-8 of a string in one step when the input is
contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link the BON8 functions from the other binary format pages
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the bulk scan flag after the input, not BON8
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BSON keys in bulk from contiguous input
BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures of the bulk-read tests
- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
when exceptions are disabled: they catch the parse errors of invalid
input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the explicit basic_json instantiation into its own test file
Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).
Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Convert the bytes of the BON8 test strings explicitly
The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests fake a container size of UINT32_MAX + 1, which does not fit
into a 32-bit std::size_t: MSVC rejects the truncation (C4305/C4309
with /WX), and clang-cl wraps the size to 0 so nothing throws. Guard
them with SIZE_MAX > UINT32_MAX like the tests from #5584.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When using cbor_tag_handler_t::store, tags 0xD8-0xDB previously assumed
that the tagged item was a byte string, unconditionally attempting to
parse binary data and failing on valid CBOR documents containing tags
applied to integers, strings, arrays, or objects (such as self-describe
tag 55799).
Check whether the tagged data item is a byte string (0x40-0x5B or 0x5F).
If it is a byte string, store the subtype on the binary value as before.
Otherwise, iteratively process the tagged value in the driver loop using
item_read so that chained tags do not consume native stack space.
Part of #5316.
Signed-off-by: ReturnKartikey <kartikeynegi2000.work@gmail.com>
* Cut test suite runtime in binary roundtrips and integer sweeps
The Linux CI jobs pass --no-skip, so skip() does not help there.
Parse each corpus file once in the binary roundtrip loops instead of
four times. Sample the 16-bit integer ranges with stride 7 (still hits
every low byte) and always keep the endpoints.
Also drop the 5M-node parse test to 500k, which still covers the
non-recursive destructor, and move jeopardy.json into its own skipped
test so the cheaper binary-format size checks actually run.
See #5418.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Drop useless int32_t casts in the sampled integer loops
ci_test_gcc compiles with -Werror=useless-cast. On that compiler
int32_t is int, so static_cast<int32_t> of the loop bound is an
error. The bounds are already int, and the sampled values do not
change.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Revert unit-binary_formats.cpp to develop and fix comment
Revert tests/src/unit-binary_formats.cpp to its develop state.
The test-case split made valgrind jobs slower instead of faster,
because the cheaper corpus files (canada/twitter/citm/sample)
now ran under valgrind where they never did before.
Fix the next_integer_sample comment: the function has no 'first'
parameter, so describe what the function actually does.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Throw instead of writing MessagePack lengths beyond UINT32_MAX
MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.
The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the MessagePack length check
- mark to_msgpack_length's value as used when exceptions are disabled
(-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
documentation style check requires
- create the test's string value from its type: constructing it from a
beyond_uint32_string_t considers the std::filesystem::path conversion,
which libstdc++ 10 reports as ambiguous for a class derived from
std::string (clang 13)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the MessagePack string length test for clang with libstdc++ 10
C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the MessagePack string test type and its alias in one block
astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>