Commit Graph

1025 Commits

Author SHA1 Message Date
Niels Lohmann
0a97d94497 Fix CI: useless cast in the Zmij digit writer and snprintf truncation
- ci_test_gcc (Linux x86-64): static_cast<std::size_t>(d.significand % 100)
  was a useless cast (a std::uint64_t prvalue, the same type as
  std::size_t there); cast a named variable instead.
- ci_test_gcc: -Werror=format-truncation for snprintf("%.*e") in
  unit-to_chars.cpp, whose precision GCC cannot bound; write the
  neighboring decimal with a stream (classic locale, std::scientific),
  which gives the same text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:19 +02:00
Niels Lohmann
bee66ec810 Write doubles with the shortest digits (Zmij)
dump() writes doubles with the conversion of Zmij by Victor Zverovich
(https://github.com/vitaut/zmij, MIT), ported to C++11 in
detail/conversions/zmij.hpp: the shortest decimal in the rounding
interval, the closest one if there are several. Grisu2 does not always
find the shortest digits; about 0.14% of random doubles are now written
differently (0.08% with fewer digits, 0.06% with the closest last
digit); short decimals such as 0.1 or 2555.56 are not affected. float
keeps Grisu2.

The layout of doubles is unchanged, but written differently: the digits
are converted eight at a time (the BCD conversion of Xiang JunBo, as in
Zmij) and stored with one byte swap per eight digits; leading and
trailing zeros are counted from those bytes; and the layouts of
format_buffer() are written with fixed-size moves instead of per-digit
loops and moves of the buffer (to_chars() uses a local buffer if the
caller's is shorter than the 41 bytes this may write).

The powers of ten come from the table for number parsing, adjusted
where it holds them rounded up, and from the compressed tables of Zmij
beyond 10^308. json::dump() gets faster on floats: canada -53%,
numbers -46%, mesh -37%, marine_ik -30%.

Tests: the powers of ten recomputed with a small big-integer; for random
doubles, all powers of two and of ten and their neighbors, and boundary
values: the output reads back as the same value, no decimal with one
digit fewer does, the layout equals that of format_buffer() for the same
digits, and (C++17) the digits equal those of std::to_chars.
The size ratios of canada.json in unit-binary_formats.cpp and one
expectation in unit-to_chars.cpp change with the shorter output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:18 +02:00
Niels Lohmann
1df24a0bd9 Write compact dumps of json_view without a library call per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer: the same walk and output, with the write position in a local
variable (stores through char pointers would otherwise force a reload
of the buffer's members after each one), strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left (the buffer keeps 64 bytes of slack), and decoded
strings copied in runs up to the next quote, backslash, or control
character. Documents that are not edited are walked through the node
array in order, so that a frame only needs the end of its container,
and integer tokens are read from the source directly. The innermost
open container is kept in local variables, and the stack holds only
the ones around it; the stack starts in a local array of 32 and moves
to the heap only for deeper nesting (its address does not escape, so
its pointers stay in registers). Dumps of shallow documents thus
allocate only the output, whose first size includes the slack, so it
does not grow just before the end.

The long copies are out of line: otherwise, the compiler merges the
fixed-size moves into the same library call.

These techniques come from the prototype; the writer lost them when the
view was split into pull requests, which made dump() 2 to 3 times
slower.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:17 +02:00
Niels Lohmann
779ae7fffc Address the cpplint findings of json_document images
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:14 +02:00
Niels Lohmann
8db804ca8f Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).

- save() writes an edited document in its current state, in document
  order (floats that are not finite become null, as in dump()); the
  same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
  load(vector&&) keeps it without a copy. The nodes are copied (aligned,
  and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
  bounds, UTF-8, strings of the source, number tokens and their values);
  bounds checks structure and bounds, so that reading and serializing
  stay safe; none trusts the image.

A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.

As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.

The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.

Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:13 +02:00
Niels Lohmann
dcc81f43f0 Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
  number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names

The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.

Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:09 +02:00
Niels Lohmann
383ce0b040 Cover the edit storage in tests
Test edits of an empty document and assignments through a view of a
value that is no longer part of the document; copy the entries of a
block through deref() (one path for links and values); mark the
4 GiB limit and the returns of find_parent() that no document reaches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:06 +02:00
Niels Lohmann
51ee239b3c Address the clang-tidy findings of the editable documents
Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:04 +02:00
Niels Lohmann
6e50766dc3 Add editable documents with set() and push_back()
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:

- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
  object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
  the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)

Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.

Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.

Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:03 +02:00
Niels Lohmann
d4fab0971f Walk the index of json_view through a navigation policy
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:

- node flags `edited` (a string or number token in the edit arena),
  `moved` (the elements of an array/object live in a separate sequence),
  and `is_new` (no source position), and link nodes (kind_link) that
  stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
  values, and the edit arena

materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:02 +02:00
Niels Lohmann
eede67ca92 Fix old clang: do not declare the defaulted document_data() noexcept
With the nested struct object_index, clang 4 (and, by the same bug, the
clang 3.x of ci_test_compilers_clang) rejects the explicitly noexcept
defaulted constructor: "default member initializer for 'indexes' needed
within definition of enclosing class 'document_data' outside of member
functions". Nothing depends on the constructor being noexcept, so let it
take the implicit exception specification.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:01 +02:00
Niels Lohmann
82b31f31b8 Fix CI: useless casts of the key hash of the view's object index
GCC -Werror=useless-cast on Linux x86-64 rejects
static_cast<std::size_t>(key_hash(...)): the call returns a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Store the hash in a variable
and cast that, which GCC does not report.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:00 +02:00
Niels Lohmann
904c8c6710 Index the large objects of json_view
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.

Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.

Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:58 +02:00
Niels Lohmann
1f0c3be6f3 Address the clang-tidy findings of the SIMD scan
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:56 +02:00
Niels Lohmann
c3b51addbf Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:55 +02:00
Niels Lohmann
ecf9df7c45 Convert the floats of json_view from the digit layout
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For floats and doubles with at most 19
digits, the value is now read from that layout: the digits eight at a time,
without scanning the token, and rounded by the library's conversion core
(detail::decimal_to_float(): Clinger's fast path where both operands are
exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19
digits). It rounds correctly, so the values are those of parse(); other
tokens and types keep the library's conversion of the whole token.

get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s.

Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and
those of float) to the bit-for-bit comparison with parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:51 +02:00
Niels Lohmann
a7fa8d04e8 Address the cpplint findings of json_view's comparisons
compare.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:49 +02:00
Niels Lohmann
677507137b Address the clang-tidy findings of the comparisons
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:49 +02:00
Niels Lohmann
e842f3a68f Add comparisons to json_view
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.

Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:47 +02:00
Niels Lohmann
bfb2b0cb48 Mark the cases of the view's serializer that tests cannot reach
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:45 +02:00
Niels Lohmann
8cccae029b Address the clang-tidy findings of dump()
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:44 +02:00
Niels Lohmann
83d72fd7a7 Add dump() to json_view
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.

The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.

Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:42 +02:00
Niels Lohmann
da971a52c8 Fix CI: useless cast in the array index check of the view's JSON pointers
GCC -Werror=useless-cast on Linux x86-64 rejected
static_cast<std::uint64_t>((std::numeric_limits<std::size_t>::max)()),
as both are the same type there. Compare without the cast: std::size_t
converts to std::uint64_t implicitly on every platform.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:41 +02:00
Niels Lohmann
a7256deba3 Inline get() of arithmetic values of json_view
get<T>() of arithmetic types is inlined down to the conversion, so that
its checks of the node kind merge with those of the caller, and reading
an integer needs no call. Traversing every value: citm_catalog -6%,
marine_ik -5%, numbers and twitter -3%, mesh -2.5%, canada -1% (and
more above the float conversion from the digit layout: citm_catalog
-14%, marine_ik -11%).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:39 +02:00
Niels Lohmann
3a4eddf4ac Address the clang-tidy findings of values and JSON pointers
get_string() and number_token() return braced lists; the test compares
floats by their bit patterns instead of with memcmp, uses std::any_of, and
marks a fixed seed, a default member initializer (needed by GCC's
-Weffc++), and a string search.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:38 +02:00
Niels Lohmann
0650a48659 Add values and JSON pointers to json_view
basic_json_view gains get<T>(), get_to(), value() with keys and JSON
pointers, and operator[], at(), and contains() with JSON pointers, plus
two functions basic_json has no counterpart for:

- get_string(): the string without a copy (a string_view into the source,
  or into the decoded strings for strings with escapes)
- number_token(): the text of a number as it appears in the source

get<T>() converts arithmetic types, strings (also string_view_t),
std::nullptr_t, std::vector, maps with string keys, and views directly;
floats are converted from the digit layout recorded by the parser with the
library's conversion chain, so the values are bit-identical to parse().
Other types, including user types with from_json(), go through
materialize().

The exceptions are those of basic_json, message included. Where const
basic_json has undefined behavior (a missing key or an index out of range
with operator[] and a JSON pointer), the result is a discarded view;
value() returns the default wherever basic_json catches out_of_range, and
contains() never throws. Array indices of JSON pointers follow
json_pointer's rules (parse_error.106/109, out_of_range.404/410).

Tests compare the conversions of 2,000 generated documents, 20,000 float
tokens (double and float, bit for bit), and every JSON pointer of 1,000
documents with basic_json, and the exceptions for malformed pointers.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:36 +02:00
Niels Lohmann
29bb5c48b8 Give code outside basic_json the reference tokens of a json_pointer
detail::json_pointer_access returns the reference tokens of a pointer, so
that code resolving pointers without a basic_json value (such as the
zero-copy view) does not have to parse to_string() again. No change in
behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:35 +02:00
Niels Lohmann
c2d7177d30 Address the clang-tidy findings of element access and iteration
Marks the default initializer of the item's index string (needed by GCC's
-Weffc++) and, in the test, an escaped literal and a comparison of find()
with end(), which is what the test is about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:33 +02:00
Niels Lohmann
7dae258350 Add element access, lookup, and iteration to json_view
basic_json_view gains the read-only access functions of basic_json:
operator[] and at() with keys and indices, front(), back(), find(),
contains(), count(), begin()/end(), items() (with structured bindings from
C++17 on), and type_name(). They throw the exceptions (ids and messages)
that the const functions of basic_json throw; where basic_json has
undefined behavior (operator[] with a missing key or an index out of
range, front()/back() of an empty container), the view returns a
discarded view or throws invalid_iterator.214.

Objects are iterated in document order, and all members are visited. With
duplicate keys, lookups find the first member, so that a lookup can stop
at the first match; parse() keeps the last value. Keys of up to 16 bytes
are compared with two overlapping loads instead of memcmp, and most keys
are rejected by their length alone, from the index.

The iterators and items live in detail/view/iterator.hpp, the lookups in
detail/view/lookup.hpp. Tests compare every element and member of 2,000
generated documents with ordered_json, keys of every length around the
load sizes, the exceptions against const basic_json, and the iterators.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:32 +02:00
Niels Lohmann
7da943c5c4 Name JSON types without a basic_json value
basic_json::type_name() now calls detail::value_type_name(value_t), so
that code which reports types without a basic_json value at hand, such as
the zero-copy view, uses the same names in its exception messages. No
change in behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:31 +02:00
Niels Lohmann
20d0723b67 Address the cpplint findings of json_view's materialize()
materialize.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:58:48 +02:00
Niels Lohmann
06bebe2af6 Mark the code of json_view that tests cannot reach
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:58:47 +02:00
Niels Lohmann
723cf14be4 Address the clang-tidy findings of json_document and json_view
- the input dispatch takes byte ranges by const reference and reads the
  size once (which also settles a finding of the static analyzer); input
  adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
  constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
  the other tests

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:58:46 +02:00
Niels Lohmann
cf352d4ef5 Add json_document and json_view: parse, accept, types, materialize
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:

- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
  inputs, owning rvalue strings, streams, and other inputs), parse_copy,
  accept, read, root, is_discarded, source, owns_source, node_count,
  memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
  materialize (the value parse() would produce, built by the same SAX
  handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
  ordered_json_view

A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.

The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:57:43 +02:00
Niels Lohmann
b88e5f9107 Keep the behavior-changing configuration readable after json.hpp
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann
d23803fd32 Decode \u escapes with a table in the lexer
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.

json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.

Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:12 +02:00
Niels Lohmann
e91fdad877 Find the stop byte of a string run without a byte loop
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.

The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:12 +02:00
Niels Lohmann
9c71689715 Convert long doubles under a multi-byte decimal point completely
The strtold fallback, which is left only for long double formats that
are not binary64 (x87, binary128), substituted the first byte of the
locale's decimal point for '.'. Under a locale whose decimal point is
longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B),
strtold stopped there and the value was truncated at the decimal point.
A longer decimal point is now put into a copy of the token.

The test "locale with a multi-byte decimal point" now compares the long
double values with those of the "C" locale; with x87 long doubles it
failed before.

Fixes #5660.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:11 +02:00
Niels Lohmann
44ec53c77b Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:

- The token is split into sign, significand w (at most 19 digits), and
  decimal exponent q, using the positions of the decimal point and the
  exponent that the scanners already recorded, so no character is
  classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
  an exact big-integer comparison with the midpoint between the two
  candidates (the digit comparison of fast_float, simplified).

This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.

convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.

Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.

Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:11 +02:00
Niels Lohmann
6a073dbae4 Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().

Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.

Fixes #5652.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:34 +02:00
Niels Lohmann
fdcc569eee Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.

Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.

Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).

Fixes #5666.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:31 +02:00
Niels Lohmann
6fc0d501f3 Move the user-defined string literals to <nlohmann/json_literals.hpp> and add JSON_NO_AUTOMATIC_UDLS (#5610)
* Add JSON_NO_UDLS to leave out the user-defined string literals

The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mention JSON_NO_UDLS in the list of exported module symbols

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the user-defined string literals to <nlohmann/json_literals.hpp>

Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.

The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test

- Suppress clang-tidy misc-header-include-cycle on the intentional mutual
  include of json.hpp and json_literals.hpp.
- Use operator"" _json with a space for GCC 4.8 in the test's detection
  aliases, as the header does.
- Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Declare the literal operators through a local macro

The GCC 4.8 spacing condition was repeated for both operator definitions
and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix)
now selects operator""##suffix or operator"" suffix in one place and is
undefined at the end of json_literals.hpp.

Suggested by gregmarr in review.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:17 +02:00
Niels Lohmann
67435c9c7e Give a deep copy its type only after its container exists (#5721)
When copying a value nested deeper than 128 levels, and an allocation
fails while an inner array or object is being copied, the partially
built copy ended up with an element typed array/object but holding a
null pointer. That element was already a fully constructed member of
its parent's container, so destroying the parent during stack
unwinding dereferenced the null pointer (release builds) or failed
assert_invariant() (debug builds), instead of letting std::bad_alloc
reach the caller.

copy_iteratively() set a pending worklist element's type right after
popping it, before the next loop iteration created its container in
copy_array_level()/copy_object_level(). Move that type assignment into
those two functions, right after the container is successfully
created, and drop the premature one in copy_iteratively(), so a
half-built element stays a null value - as copy_shallow()'s comment
already promised - until it can safely hold one.

Add a regression test to tests/src/unit-allocator.cpp that copies a
value nested 130 levels deep (both arrays and objects, with a
std::map- and an ordered_map-backed object_t) and fails every
allocation of the copy in turn: each attempt must throw std::bad_alloc
without crashing, and the source must stay unchanged.

Fixes #5640.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:00 +02:00
Niels Lohmann
2ea6d8c127 Require the found key to equal the looked-up key when comparing objects (#5720)
For an object type whose comparator treats unequal keys as equivalent
(for example a std::map with a case-insensitive comparator),
compare_iteratively() looked up a mismatched left key in the right
object with find(), which uses the object's own comparator, and
accepted whatever entry it found without checking that the keys are
actually equal. A case-insensitive comparator then found "KEY" for
"key", so two objects nested past the recursion bound (or at every
depth with JSON_NO_THREAD_LOCAL) could compare equal even though the
object type's own operator== - and basic_json itself, below the bound
- consider them different.

Accept the found entry only if its key equals (not just compares
equivalent to) the looked-up key.

Fixes #5655.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:56 +02:00
Niels Lohmann
66877675b1 Check the iterator range for binary values in basic_json(first, last) (#5719)
basic_json(first, last) treated value_t::binary like the structured
types (array, object) in the range check, so it always copied the
whole binary value regardless of the iterators, even for an empty
range such as (b.end(), b.end()). The other primitive types (number,
boolean, string) already reject such a range with
invalid_iterator.204, and erase(first, last) already does the same
for binary values, so this made the constructor inconsistent with
both. Move case value_t::binary into the group of checked primitive
types.

Also update the two matching passages in basic_json.md that describe
overload 7, and add a version-history note.

Fixes #5670.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:52 +02:00
Niels Lohmann
7fd6895788 Hide a discarded container's content from the parser callback (#5706)
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.

The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.

Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.

Fixes #5643.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:48 +02:00
Niels Lohmann
6ae17630a4 Reject integral keys for contains(), find(), and count() at compile time (#5705)
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is
a null pointer constant, so it converts to a null const char*, and the
overloads taking const typename object_t::key_type& accepted it by
constructing a std::string from that null pointer, which is undefined
behavior (a crash with both libc++ and libstdc++). value(0, default_value)
had the same problem in C++11, where the object comparator is not
transparent.

Add deleted overloads for integral arguments to contains(), find()
(const and non-const), count(), and value() so that these calls are
compile errors in every supported language mode instead of crashing.
Calls with string, string_view, json_pointer, and size-typed element
access (at(), operator[](), erase()) are unaffected.

Fixes #5657.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:45 +02:00
Niels Lohmann
5bd766aa50 Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.

calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.

Fixes #5675.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:41 +02:00
Niels Lohmann
4bb1b14b06 Trim the compiler-appended NUL from wide/UTF string literals too (#5702)
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.

Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.

Fixes #5658.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:37 +02:00
Niels Lohmann
bfe0f32d71 Fix std::terminate and null pointer access in input_stream_adapter (#5699)
Parsing from a std::istream crashed in two unusual but valid stream
states, both in input_stream_adapter:

- With eofbit in the stream's exceptions() mask, get_character() sets
  eofbit via is->clear(), which throws std::ios_base::failure. While
  that exception unwinds, ~input_stream_adapter() called clear() again
  to reset eofbit, which is still set and still in the exception mask,
  so it throws a second time out of the (implicitly noexcept)
  destructor and std::terminate() is called. The destructor now only
  calls clear() if a bit other than eofbit remains set, so the first
  exception can propagate normally.
- For an std::istream without a stream buffer (rdbuf() == nullptr,
  e.g. std::istream(nullptr)), the constructor stored the null
  pointer without checking it, and get_character() dereferenced it.
  input_adapter(std::istream&) now throws parse_error.101 for such a
  stream, the same as it already does for a null FILE* or char*.

Added regression tests to unit-deserialization.cpp and, for the
JSON_PRECISE_STREAM_POSITION variant of get_character(), to
unit-precise-stream-position.cpp; both crashed before this fix.
Documented the two exceptions in parse.md and operator_gtgt.md.

Fixes #5646.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-30 20:07:33 +02:00