Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:
- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)
Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.
Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.
Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.
Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.
Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.
Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.
json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.
Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A workflow runs compare.py with pinned downloads on GitHub-hosted
Ubuntu runners and shows the results as the job summary and as an
artifact: started by hand (workflow_dispatch: x86-64 or AArch64, GCC or
Clang), or when a pull request gets the label "benchmark" (both
architectures, GCC). The label trigger gives numbers before the
workflow is on the default branch, which workflow_dispatch needs.
Shared runners are noisy, so the numbers show where json_view stands on
another architecture; published numbers still need a quiet machine.
compare.py takes the CPU name from lscpu where /proc/cpuinfo has none
(AArch64 Linux), and falls back to the architecture. Checked in Linux
containers (AArch64, Clang 15 and GCC 9, offline with the pinned
archives).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
compare.py --download now checks the SHA-256 of the yyjson 0.13.0,
simdjson 4.6.11, and Boost 1.92.0 archives. It unpacks each archive
once (Boost's directory is boost_1_92_0) and, where Python supports it,
with the 'data' filter.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
tests/benchmarks/json_view/ holds the comparison with other libraries,
which is not built by CMake or run by CI:
- bench_view.cpp: parse, traverse, select, and dump of twitter,
citm_catalog, canada, jeopardy, a single tweet, and a JSON-RPC request,
with json_view, yyjson, simdjson (DOM and On-Demand), Boost.JSON, and
json::parse; all engines must agree on every document before anything
is timed, and run interleaved in every round
- bench_corpus.cpp: parse, traverse, and dump of any list of files
- compare.py: builds both against include/ with the libraries of the
system (or pinned downloads), runs them, and writes the results with
what is needed to reproduce them (date, commit, CPU, OS, compiler,
flags, library versions) to results/<date>-<host>.md and .csv; only the
Python 3 standard library is used
- README.md: how to run it, what is measured, and which features the
engines have, so the numbers can be read correctly
Boost.JSON is optional (JSON_VIEW_BENCH_BOOST).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For floats and doubles with at most 19
digits, the value is now read from that layout: the digits eight at a time,
without scanning the token, and rounded by the library's conversion core
(detail::decimal_to_float(): Clinger's fast path where both operands are
exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19
digits). It rounds correctly, so the values are those of parse(); other
tokens and types keep the library's conversion of the whole token.
get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s.
Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and
those of float) to the bit-for-bit comparison with parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.
Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.
The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.
Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: exception_of() and without_path() exist only
with exceptions (they catch outside a CHECK_THROWS, which aborts with
JSON_NOEXCEPTION); compile the comparisons of the conversion, value(),
and JSON pointer errors only with exceptions as well.
- ci_test_gcc: -Werror=unused-result for static_cast<void>(j.contains(p))
(GCC's warn_unused_result ignores a cast to void); store the result.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_string() and number_token() return braced lists; the test compares
floats by their bit patterns instead of with memcmp, uses std::any_of, and
marks a fixed seed, a default member initializer (needed by GCC's
-Weffc++), and a string search.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains get<T>(), get_to(), value() with keys and JSON
pointers, and operator[], at(), and contains() with JSON pointers, plus
two functions basic_json has no counterpart for:
- get_string(): the string without a copy (a string_view into the source,
or into the decoded strings for strings with escapes)
- number_token(): the text of a number as it appears in the source
get<T>() converts arithmetic types, strings (also string_view_t),
std::nullptr_t, std::vector, maps with string keys, and views directly;
floats are converted from the digit layout recorded by the parser with the
library's conversion chain, so the values are bit-identical to parse().
Other types, including user types with from_json(), go through
materialize().
The exceptions are those of basic_json, message included. Where const
basic_json has undefined behavior (a missing key or an index out of range
with operator[] and a JSON pointer), the result is a discarded view;
value() returns the default wherever basic_json catches out_of_range, and
contains() never throws. Array indices of JSON pointers follow
json_pointer's rules (parse_error.106/109, out_of_range.404/410).
Tests compare the conversions of 2,000 generated documents, 20,000 float
tokens (double and float, bit for bit), and every JSON pointer of 1,000
documents with basic_json, and the exceptions for malformed pointers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: the element access tests compare the exceptions
of json_view and basic_json through exception_of(), which catches them
outside a CHECK_THROWS; with JSON_NOEXCEPTION the first one aborted the
test. Compile those comparisons only with exceptions.
- clang 3.6: value-initialize a const json_view, as in the tests of
json-view/10-view-document.
- Format three new documentation examples with the pinned astyle, which
the "check" job runs once it gets past the amalgamation step.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Marks the default initializer of the item's index string (needed by GCC's
-Weffc++) and, in the test, an escaped literal and a comparison of find()
with end(), which is what the test is about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains the read-only access functions of basic_json:
operator[] and at() with keys and indices, front(), back(), find(),
contains(), count(), begin()/end(), items() (with structured bindings from
C++17 on), and type_name(). They throw the exceptions (ids and messages)
that the const functions of basic_json throw; where basic_json has
undefined behavior (operator[] with a missing key or an index out of
range, front()/back() of an empty container), the view returns a
discarded view or throws invalid_iterator.214.
Objects are iterated in document order, and all members are visited. With
duplicate keys, lookups find the first member, so that a lookup can stop
at the first match; parse() keeps the last value. Keys of up to 16 bytes
are compared with two overlapping loads instead of memcmp, and most keys
are rejected by their length alone, from the index.
The iterators and items live in detail/view/iterator.hpp, the lookups in
detail/view/lookup.hpp. Tests compare every element and member of 2,000
generated documents with ordered_json, keys of every length around the
load sizes, the exceptions against const basic_json, and the iterators.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: the helpers that compare the exceptions of
json_document::parse() and json::parse() catch them outside a
CHECK_THROWS, so with JSON_NOEXCEPTION the first parse error aborted
the test. Compile those comparisons only with exceptions, as
unit-class_parser.cpp does.
- ci_test_gcc: -Werror=unused-result for CHECK_THROWS_AS(json_document::
parse(...)); assign the result to a dummy document.
- ci_test_compilers_clang (3.6): `const json_view invalid;` needs a
user-provided default constructor there (CWG 253); value-initialize it.
- ci_test_single_header: json_view.hpp now exists as a single header and
contains the internal view headers, so unit-json_view_builder.cpp
includes it instead of the detail headers in that mode, and the test
is built again with the single header.
- Regenerate single_include/nlohmann/json_view.hpp for the builder change
merged from json-view/08-view-builder.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- the input dispatch takes byte ranges by const reference and reads the
size once (which also settles a finding of the static analyzer); input
adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
benchmarks_view.cpp adds ViewParse, ViewRead (a reused document),
ViewParseIndented, ViewAccept, and ViewMaterialize on the files of
ParseString, so that each row can be read against the json::parse row of
the same file; benchmarks.cpp gains Accept (json::accept) as the
counterpart of ViewAccept. The view benchmarks are built only if the
header directory has json_view.hpp, so that older versions can still be
benchmarked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
fuzzer-parse_json_view.cpp checks for every input that
json_document::accept agrees with json::accept, that an accepted input
materializes to the value json::parse returns, and that a rejected input
makes both throw the same exception with the same message. It is built
like the other fuzzers (tests/Makefile, and the root Makefile's
fuzz_testing_json_view target, which starts from the JSON test corpus) and
listed in tests/fuzzing.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The nlohmann.json module includes json_view.hpp in its global module
fragment and exports basic_json_document, basic_json_view, and the four
aliases next to basic_json, json, and ordered_json. features/modules.md
lists them, and tests/module_cpp20 parses a document, so that a missing
export fails the ci_module_cpp20 job.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
unit-json_view.cpp: type queries, size, and empty against basic_json;
materialize() against parse() (json and ordered_json, generated documents,
duplicate keys, 100,000 levels of nesting, parent pointers with
JSON_DIAGNOSTICS); parse errors and their messages equal to parse() for
malformed inputs and all option combinations; the overflow of a float
document (1e39, 3.4028236e38) as in parse(); NUL and BOM; borrowed and
owned inputs (strings, C strings, literals, vectors, string_view, streams,
wide strings, parse_copy, and iterator ranges over pointers, vectors,
strings, and lists); reuse with read(); moves; shrink_to_fit() of the index
and of the decoded strings; source offsets.
unit-json_view_macros.cpp includes the header without
JSON_TEST_KEEP_MACROS, as users do: the view must not depend on the macros
json.hpp undefines, and must not leak its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
constant) for `TrailingCommas && cur() == ']'` and the like when the
option is off. Route the template arguments through a static enabled()
function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
single_include/, which does not contain the internal
nlohmann/detail/view headers. Build that test only with the multiple
headers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Numbers of more than 19 digits at the boundary of the largest double are
decided by the exact comparison with the midpoint in the overflow check.
The check uses the floating-point type of the document: with float, the
view rejects what parse() rejects (1e39, 3.4028236e38, the midpoint between
the largest float and 2^128), and double documents are not affected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- tables as std::array; the frames of the first 64 levels stay a C array
(not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
second table
- the parse failure is private, with an accessor; the special member
functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.
Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.
Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.
json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.
Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast rejected static_cast<std::size_t>(next() % n):
on 64-bit Linux std::uint64_t and std::size_t are the same type, while
the cast is needed where std::size_t is 32 bits wide. Draw the sizes from
a 32-bit value instead, which converts to std::size_t implicitly on every
platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.
The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:
- The token is split into sign, significand w (at most 19 digits), and
decimal exponent q, using the positions of the decimal point and the
exponent that the scanners already recorded, so no character is
classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
an exact big-integer comparison with the midpoint between the two
candidates (the digit comparison of fast_float, simplified).
This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.
convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.
Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.
Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...))
because parse() is [[nodiscard]]; assign the result to a dummy json as the
other tests do. clang-tidy flagged longer.find('.') == npos with
abseil-string-find-str-contains; store the position in a variable first.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).
The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.
json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:
- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
calls whose result was discarded. Check the result, like the other
resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
callbacks, which cannot throw but were not declared noexcept.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Look up the locale decimal point at conversion time, not lexer construction
The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.
token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.
Fixes#5198
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop the strtod retry loop when the decimal point is unchanged
convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.
Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Weffc++ errors in the #5198 locale test
GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the key type when rejecting non-string CBOR/MessagePack map keys
CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:
syntax error while parsing MessagePack object key: only string keys
are supported, but found nil; last byte: 0xC0
The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.
Refs #2766, #3381
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point the MessagePack key note to the spec's profile section
The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).
Write into a caller-owned vector instead, so nothing is reserved.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests added by #5515 fail two ways on develop:
- clang-tidy reports the size() overrides of huge_string and huge_binary
(readability-convert-member-functions-to-static) and the non-const
test value (misc-const-correctness); mark them like the #5584 types
- clang with libstdc++ 10 cannot compile the file for C++17: the
std::filesystem::path conversion considered for huge_string, a class
derived from std::string, is ambiguous. Guard it with
JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same
reason, and define that macro before both test blocks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add BON8 support
Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.
The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.
The round-trip tests need the .bon8 files of json_test_data 3.2.0.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Address review comments
- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Select the BON8 float prefix by type
get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Rename a test variable that Flawfinder mistakes for read()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures
- compare the float in write_bon8_float with number_float_t constants,
so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BON8 strings in bulk from contiguous input
- copy the valid UTF-8 of a string in one step when the input is
contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link the BON8 functions from the other binary format pages
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the bulk scan flag after the input, not BON8
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BSON keys in bulk from contiguous input
BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures of the bulk-read tests
- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
when exceptions are disabled: they catch the parse errors of invalid
input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the explicit basic_json instantiation into its own test file
Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).
Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Convert the bytes of the BON8 test strings explicitly
The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests fake a container size of UINT32_MAX + 1, which does not fit
into a 32-bit std::size_t: MSVC rejects the truncation (C4305/C4309
with /WX), and clang-cl wraps the size to 0 so nothing throws. Guard
them with SIZE_MAX > UINT32_MAX like the tests from #5584.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When using cbor_tag_handler_t::store, tags 0xD8-0xDB previously assumed
that the tagged item was a byte string, unconditionally attempting to
parse binary data and failing on valid CBOR documents containing tags
applied to integers, strings, arrays, or objects (such as self-describe
tag 55799).
Check whether the tagged data item is a byte string (0x40-0x5B or 0x5F).
If it is a byte string, store the subtype on the binary value as before.
Otherwise, iteratively process the tagged value in the driver loop using
item_read so that chained tags do not consume native stack space.
Part of #5316.
Signed-off-by: ReturnKartikey <kartikeynegi2000.work@gmail.com>
* Cut test suite runtime in binary roundtrips and integer sweeps
The Linux CI jobs pass --no-skip, so skip() does not help there.
Parse each corpus file once in the binary roundtrip loops instead of
four times. Sample the 16-bit integer ranges with stride 7 (still hits
every low byte) and always keep the endpoints.
Also drop the 5M-node parse test to 500k, which still covers the
non-recursive destructor, and move jeopardy.json into its own skipped
test so the cheaper binary-format size checks actually run.
See #5418.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Drop useless int32_t casts in the sampled integer loops
ci_test_gcc compiles with -Werror=useless-cast. On that compiler
int32_t is int, so static_cast<int32_t> of the loop bound is an
error. The bounds are already int, and the sampled values do not
change.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Revert unit-binary_formats.cpp to develop and fix comment
Revert tests/src/unit-binary_formats.cpp to its develop state.
The test-case split made valgrind jobs slower instead of faster,
because the cheaper corpus files (canada/twitter/citm/sample)
now ran under valgrind where they never did before.
Fix the next_integer_sample comment: the function has no 'first'
parameter, so describe what the function actually does.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Throw instead of writing MessagePack lengths beyond UINT32_MAX
MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.
The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the MessagePack length check
- mark to_msgpack_length's value as used when exceptions are disabled
(-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
documentation style check requires
- create the test's string value from its type: constructing it from a
beyond_uint32_string_t considers the std::filesystem::path conversion,
which libstdc++ 10 reports as ambiguous for a class derived from
std::string (clang 13)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the MessagePack string length test for clang with libstdc++ 10
C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the MessagePack string test type and its alias in one block
astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>