Commit Graph

311 Commits

Author SHA1 Message Date
Niels Lohmann
a32f61eb98 Fix update() and merge_patch() when the argument is *this or one of its members (#5678) 2026-09-30 23:04:57 +02:00
Niels Lohmann
1d675cdb46 Fix CI jobs that check less than they claim; move arm64 to GitHub (#5733)
* Fix the ci_cmake_flags wiring so every option is checked

The CMake 3.31.6 flag list referred to itself before it was defined,
so only JSON_BuildTests was checked with that version. The targets for
the CMake running the build ("_2") were created but never added to
ci_cmake_flags, and the three versions shared one build directory.
JSON_StrictNulHandling was not in the list at all.

Use the 3.5.0 list for 3.31.6, add JSON_StrictNulHandling, and create
one ci_cmake_flag_<flag> target per option for the running CMake with
its own build directory. Also use the function parameter in the
COMMENT, refresh the stale version comment, and let ci_clean remove
the downloaded cmake-<version> directories instead of the long-gone
cmake-3.5.0-Darwin64.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix ci_test_clang_libcxx_cxx* jobs silently building without warnings

CMake only seeds CMAKE_CXX_FLAGS from the CXXFLAGS environment variable
when the cache entry is unset, so the explicit -DCMAKE_CXX_FLAGS="-stdlib=libc++"
argument made it ignore CXXFLAGS="${CLANG_CXXFLAGS}" entirely. The six
ci_test_standards_clang (..., libcxx) jobs therefore compiled without
-Weverything/-Werror while their libstdc++ siblings did use them.

Pass -stdlib=libc++ through the same CXXFLAGS value instead of a separate
-D argument, and give the target its own build directory
(build_clang_libcxx_cxx${CXX_STANDARD}) so it no longer shares a CMake
cache with the libstdc++ variant. Suppress the resulting
-Wthread-safety-negative finding from libc++'s std::mutex annotations,
which fires on doctest's reporters in this translation unit only.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the no-op AppVeyor with_win_header job

The with_win_header matrix entry patched Windows.h into
single_include/nlohmann/json.hpp before building, but JSON_MultipleHeaders
has defaulted to ON since #3532 (2022-06), so CMakeLists.txt points the
tests at include/ and the patched single header is never compiled. The
job has been a no-op VS2015 build since then.

Windows.h coverage already exists through tests/src/unit-windows_h.cpp
(#3631), which runs in every MSVC job. Delete the dead matrix entry and
its before_build steps, and cite unit-windows_h.cpp from the QA page.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove unused ci_oclint and ci_pvs_studio targets

No workflow invokes ci_oclint, ci_pvs_studio, or their tool discovery.
ci_oclint also had a side effect on every JSON_CI configure: it copied
the single header into src_single/all.cpp and added an add_executable()
for it without EXCLUDE_FROM_ALL, so a plain build compiled a 1.2 MB
translation unit that only that unused target consumed. ci_pvs_studio
duplicates the Makefile's pvs_studio target, which is kept.

Also drop the duplicate --check-level=exhaustive flag passed twice to
the same ci_cppcheck invocation.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop Dependabot from proposing astyle bumps

astyle is deliberately pinned at 3.4.13 because newer versions reformat
unrelated lines and this version defines the formatting that
check_amalgamation.yml enforces. Without an ignore rule, Dependabot
keeps opening PRs for every new astyle release (most recently #4580,
#4942, #5445, #5448), each of which fails the amalgamation check and
gets closed unmerged.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move Linux arm64 CI from dead Cirrus CI to ubuntu-24.04-arm

Cirrus CI stopped reporting check runs on develop sometime after
d10879bca (2026-05-26); every commit since has only github-actions
check runs, so .cirrus.yml silently lost its only consumer while
README.md, FILES.md and the QA page kept advertising the coverage.

Add a ci_test_arm64 job to ubuntu.yml using the same pinned
actions/checkout and lukka/get-cmake actions as the other jobs, on the
native ubuntu-24.04-arm runner, with a step that confirms uname -m
reports aarch64. Delete .cirrus.yml and its README badge and FILES.md
section, and update the QA page's arm64 row.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make scan-build fail on findings and drop irrelevant checkers (#5715 item 4a)

ci_clang_analyze ran scan-build without --status-bugs, so the job
passed whenever the ninja build succeeded, no matter what the
analyzer found ("No bugs found" in a green run gave no signal either
way). It is also missing --use-analyzer=${CLANG_TOOL}, so scan-build
picks whichever clang happens to be first on PATH inside the
silkeh/clang:dev container instead of the one this file already
selected and versioned.

Add --status-bugs and --use-analyzer=${CLANG_TOOL} to the scan-build
invocation. While here, drop the osx.*, webkit.*, fuchsia.*, and
optin.mpi.* checkers from CLANG_ANALYZER_CHECKS: none of them apply
to this portable C++ library, and leaving them enabled only adds
noise once the job can actually fail on a finding.

The job is currently clean (0 bugs), so this alone does not surface
any new finding; it only makes the existing "no bugs found" result
authoritative. This is 4a of 3 independent steps in #5715 item 4;
4b (Infer) and 4c (IWYU) still need their existing findings triaged
before --fail-on-issue/-Xiwyu --error can be added, and are handled
in separate commits.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the amalgamation/format check's file set and add BUILD.bazel (#5715 item 5)

The amalgamation/format check existed three times with three different
file sets: the Makefile's pretty/check-amalgamation, the pull_request-only
check_amalgamation.yml workflow, and the ci_test_amalgamation CMake target
that also runs on direct pushes to develop/master/release/*. The CMake
target's glob was a strict subset of the workflow's (missing the
docs/mkdocs/docs/examples/*.hpp headers, tests/abi/, tests/cmake_*/project/,
tests/cuda_example/, tests/fmt_formatter/, and tests/module_cpp20/), and it
never checked BUILD.bazel at all, so a misformatted file in any of those
paths, or a stale BUILD.bazel, could reach develop through a direct push
even though the PR-only workflow would have caught it.

Make ci_test_amalgamation glob the same roots (docs/mkdocs/docs/examples,
include, tests) and extensions (*.hpp, *.cpp, *.cu) as check_amalgamation.yml,
excluding tests/thirdparty/ and tests/abi/include/nlohmann/ the same way, and
regenerate and diff BUILD.bazel next to json.hpp/json_fwd.hpp. Also add
docs/mkdocs/docs/examples/*.hpp to the Makefile's pretty/pretty_format
targets, which were missing the four custom_*_type.hpp example headers, and
drop the stale "called by Travis" comment on check-amalgamation (Travis is
gone; nothing currently calls that Makefile target from CI).

Leaves the workflow itself untouched: it deliberately runs amalgamate.py
from a fresh develop checkout so a PR cannot change the tool that checks it.

Overlaps #5610 and #5621, which each add a new amalgamated header and touch
the same INDENT_FILES/ci_test_amalgamation/check_amalgamation.yml hunks.

Verified with `make check-amalgamation` on this branch: clean, no diff.
#5715 item 5.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Download prebuilt CMake binaries on Linux x86_64 instead of building from source (#5715 item 6)

ci_get_cmake() downloaded the source tarball of CMake 3.5.0, 3.31.6, and
4.0.0 and compiled each one completely (including CMake's own test
helpers) with -DCMAKE_POLICY_VERSION_MINIMUM=3.5 as a workaround for
building old CMake with a newer one. On CI this made the
ci_cmake_options (ci_cmake_flags) job take about 11 minutes, most of it
spent building CMake itself, even though Kitware has published
ready-to-run Linux x86_64 archives for all three of these releases
since 3.20 (lowercase platform name).

On Linux x86_64, download and unpack the prebuilt
cmake-<version>-linux-x86_64.tar.gz archive instead and point the
existing ${var} output at its bin/cmake, skipping the configure/build
steps and CMAKE_POLICY_VERSION_MINIMUM entirely. Keep the previous
source build as a fallback for any other platform (macOS, Linux
aarch64), since Kitware does not publish binaries for every
CMake/platform combination this project might build on.

Verify the downloaded archive against Kitware's own published checksum
before unpacking it: download cmake-<version>-SHA-256.txt alongside the
archive and run `sha256sum -c` on the matching line. A CI job that wgets
and untars a binary from a release page with no integrity check is a
supply-chain gap; Kitware has published this file for every release
since 3.20, so checking it costs one extra download and one grep.

As a separate, mechanical change: the ci_cmake_options job's container
only needed to stay on ubuntu:focal for the source build's
libssl-dev dependency and its own aging toolchain; now that the
Linux/x86_64 path never compiles CMake, drop libssl-dev from its apt
install line and move the job to ubuntu:24.04 (Ubuntu 20.04 left
standard support in May 2025). ci_clean already removes the
cmake-3.5.0/cmake-3.31.6/cmake-4.0.0 directories from the #5715 item 3
fix, and the prebuilt path reuses those same directory names, so no
further cleanup changes are needed.

Overlaps #5598, which edits the same ci_cmake_options matrix line in
ubuntu.yml; a rebase may be needed once that lands.

Verified locally: `cmake -S . -B build -DJSON_CI=On` configures cleanly
on macOS/arm64 (source-build fallback branch) and on Linux/x86_64 in an
ubuntu:24.04 Docker container (47 `ci_cmake_flag_*` targets generated,
one built and run successfully); `.github/workflows/ubuntu.yml` still
parses as valid YAML; downloaded the real v3.31.6 Linux x86_64 archive
and SHA-256 file from Kitware and confirmed the `grep | sha256sum -c`
pipeline both accepts the genuine file and is anchored to the exact
filename (not a prefix match).

CI must confirm: the prebuilt-binary path actually runs on the
ubuntu-latest/ubuntu:24.04 x86_64 runner, all `ci_cmake_options`
entries still pass with the new container's GCC, and the job's
runtime drops from roughly 11 minutes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make Infer fail on findings, with a type-level baseline for the ~174 pre-existing ones (#5715 item 4b)

ci_infer ran `infer run` without --fail-on-issue, so the job passed
regardless of what Pulse found; the last recorded run (35829411620,
commit 1054b2097) logged "Found 174 issues" and still went green.
report.txt was also never uploaded, so the full finding list was only
ever visible in the truncated 5-issue console excerpt.

Add a repository-root .inferconfig (auto-discovered by Infer; passing
--project-root on the `infer run` invocation makes sure it is found
even though the analysis runs from build/build_infer) that sets
fail-on-issue and disables the six PULSE issue types that made up all
174 findings in that run: PULSE_UNNECESSARY_COPY_ASSIGNMENT (129),
PULSE_UNNECESSARY_COPY (22), PULSE_UNNECESSARY_COPY_INTERMEDIATE (15),
PULSE_RESOURCE_LEAK (5), PULSE_CONST_REFABLE (2), and
PULSE_UNNECESSARY_COPY_OPTIONAL (1).

This is a deliberate, narrower fix than "triage and fix everything in
this PR": the visible sample is entirely doctest-macro copies in test
code (for example tests/src/unit-algorithms.cpp:141 and
tests/src/unit-bjdata.cpp:3706), but 169 of the 174 findings were never
uploaded anywhere and this PR cannot respectably claim to have fixed
issues it never saw, including the resource-leak and const-refable
ones that are the most likely to be genuine bugs. Disabling by issue
type is a coarser baseline than a per-finding one (Infer has no
built-in per-finding baseline short of the two-run `infer reportdiff`
workflow, which this repository does not have the CI infrastructure
for), but it has the same effect today: the job goes from always green
to green-only-when-clean-of-everything-else, so CI now fails the
moment a *new* issue type appears, and report.txt is uploaded as a
workflow artifact on every run (including failures) so the six
disabled types can be triaged and re-enabled incrementally in follow-up
PRs.

#5715 item 4b. 4a (scan-build) and 4c (IWYU) are handled in separate
commits.

Verified: .inferconfig parses as JSON, ubuntu.yml still parses as
YAML. Infer itself is not available in this environment (v1.3.0 tar.xz
requires a Linux x86_64 runner), so CI must confirm that `infer run
--project-root ... -- make` picks up .inferconfig, that fail-on-issue
takes effect, and that the six disabled types actually suppress the
existing findings without also hiding an unrelated new one.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix IWYU findings for json.hpp/json_fwd.hpp/ordered_map.hpp and make CI fail on new ones (#5715 item 4c)

ci_single_binaries ran IWYU via CMake's CXX_INCLUDE_WHAT_YOU_USE launcher
property, which only printed "Warning: include-what-you-use reported
diagnostics" without failing the build: CMake's own __run_co_compile
wrapper does not propagate the launched tool's exit code, so even
`-Xiwyu --error` could never fail `cmake --build` this way. Verified
this empirically by injecting a deliberately-unused #include and
confirming the build still exited 0.

Fix the findings from the last recorded run (issue #5715 item 4, log
35829411620):
- ordered_map.hpp: add <new> (placement new) and
  nlohmann/detail/abi_macros.hpp; drop <memory> (std::allocator is
  still visible transitively via <vector>, confirmed by full local and
  containerized test suite runs).
- json_fwd.hpp: drop <memory> (same reasoning). Keep every forward
  declaration IWYU wanted removed (adl_serializer, basic_json,
  json_pointer, ordered_map): this file's only job is to forward-declare
  them for downstream users, so "nothing in this TU uses them" is
  expected, not a real finding. Mark each with `// IWYU pragma: keep`.
- json.hpp: add <cmath>, <cstdint>, <set>, <type_traits>,
  <unordered_map>, and the detail/abi_macros.hpp, detail/input/json_sax.hpp,
  detail/meta/detected.hpp, thirdparty/hedley/hedley.hpp includes IWYU
  says it needs. Do NOT remove adl_serializer.hpp,
  detail/conversions/from_json.hpp, detail/conversions/to_json.hpp,
  detail/macro_unscope.hpp, or ordered_map.hpp as IWYU suggests: nothing
  else in include/nlohmann includes adl_serializer.hpp or
  ordered_map.hpp, so basic_json<>'s own default template arguments
  (JSONSerializer = adl_serializer, and ordered_json = basic_json<ordered_map>)
  would lose their complete type; detail/macro_unscope.hpp is what
  undoes the JSON_* macros detail/macro_scope.hpp defines earlier in
  this same file, and removing it leaks those macros into every
  translation unit that includes <nlohmann/json.hpp>. Verified by
  actually removing them in a scratch test: the header still "compiles"
  stand-alone but ordered_json and every macro-using translation unit
  break. Marked each `// IWYU pragma: keep`.

Enforce it with `iwyu_tool` (ships with IWYU, e.g. as /usr/bin/iwyu_tool
on Debian/Ubuntu) instead of relying on the launcher property: it reads
compile_commands.json (now exported project-wide under JSON_CI) and
does return a real exit code for its own analysis, independent of
CMake's wrapper. ci_single_binaries now runs it over every
src_single/*.cpp with `-Xiwyu --error`, so a *new* finding fails CI.

json.hpp itself is excluded from that hard gate: even after every fix
above, IWYU's suggestion for one remaining symbol (a container
`swap, operator!=` used somewhere via a templated comparator) is not
deterministic — repeated, otherwise-identical containerized runs
reported <set>, then <unordered_map>, then <map> as "the" header to
add/remove for the exact same source. Gating a whole CI job on a
nondeterministic suggestion would make ci_single_binaries flaky rather
than informative, so json.hpp keeps the existing informational warning
(still shown during its normal compile) without failing the build on
it. Every other one of the ~50 single-header checks is included in the
hard gate.

#5715 item 4c. 4a (scan-build) and 4b (Infer) are separate commits.

Verified: full local ctest suite (129/129) and the ci_single_binaries
target itself both green in a containerized silkeh/clang:dev run
(matching the actual CI job) after this fix; a deliberately-reintroduced
unused #include in ordered_map.hpp was confirmed to fail
`cmake --build ... --target ci_single_binaries` (exit 2) with this
change, and to pass without it, on the same container/IWYU version CI
uses. `make check-amalgamation` is clean. Compiled with Clang and GCC
at -std=c++11/14/17/20 locally with no new warnings.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:49:24 +02:00
Niels Lohmann
c261578431 Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer

The doc block of write_number() ended up above the byte_swap() helpers
added in #5286, about 80 lines from the function. It was also a plain
comment that Doxygen skips, said "write a number to output input", and
left BON8 out of the big-endian formats. Move it back onto
write_number() as a /*! block and fix the text.

write_bson() documented "@pre j.type() == value_t::object", but it
throws type_error.317 for every other type, and to_bson() relies on
that. Document the exception instead.

Explain why the CBOR binary subtype is always written with a 0xD8..0xDB
head and never in the one-byte tag form: binary_reader with
cbor_tag_handler_t::store only keeps those heads as a subtype, so
switching to write_cbor_head() would break round trips for subtypes
0..23.

Also fix the grammar of the to_char_type comment. Comments only; no
change in behavior, API or ABI.

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the duplicated UBJSON/BJData integer marker ladders

write_number_with_ubjson_prefix() (unsigned and signed overloads) and
ubjson_prefix() (number_integer and number_unsigned cases) each picked
the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their
own independent if/else ladder, and the values beyond 64 bits were
handled by a second, tag-dispatched pair of ladders. An optimized
container announces the marker of its first element via ubjson_prefix()
and then writes every element through write_number_with_ubjson_prefix(),
so the two had to be kept in lockstep by hand across four call sites.

Replace all of that with one ubjson_integer_prefix() built on
value_in_range_of<T>, and one write_ubjson_integer_payload() that
writes the value (or, for 'H', the decimal digits) for a given marker.
write_number_with_ubjson_prefix() and ubjson_prefix() keep their
signatures and now just call these two helpers.

Behavior, the public API and the ABI are unchanged. Verified with a
new regression test covering scalars and $-optimized arrays/objects at
every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for
to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing
to_ubjson/to_bjdata output before and after over the json_test_data
corpus (bit-identical).

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove dead get_char parameters in binary_reader

The non-recursive rewrite of the binary readers (#5505, #5506, #5507)
left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char
parameters dead: parse_cbor_internal() has one caller and it always
passes true, and parse_ubjson_internal() has one caller and it always
uses the true default. Both parameters, and the @param docs describing
the "reuse the last character" mode they used to select, no longer
correspond to anything.

Drop both parameters, initialise fetch/prefix unconditionally, and
update the two call sites in sax_parse(). parse_cbor_value()'s and
get_ubjson_string()'s own get_char parameters are unrelated and are
left alone; both still have a false caller.

Also delete a stray `@return whether a valid MessagePack value was
passed to the SAX parser` doxygen block that sits directly above
parse_msgpack_value()'s real doc comment, a leftover of the same
rewrite.

Behavior, the public API and the ABI are unchanged; these are private
members of detail::binary_reader. Verified by compiling with
-Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and
unit-msgpack (offline, against the stubbed test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Share the IEEE half-precision decoder between CBOR and BJData

binary_reader had two ~45-line copies of the IEEE 754 half-precision
decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is
normalised, the two blocks were identical except for the byte order
used to assemble the 16-bit half (CBOR is big endian, BJData is little
endian). Any future change to half-float decoding had to be made and
kept in sync in both places.

Add one get_half_float(format, little_endian) helper that does the two
get()/unexpect_eof() reads, assembles the half in the requested byte
order, decodes it per RFC 8949 Appendix D, and calls sax->number_float.
Both cases now just call it with their byte order; the BJData case
keeps its bjdata-only guard.

Behavior, the public API and the ABI are unchanged. Verified with a
scratch probe comparing the old and new decoders bit-for-bit (NaN by
isnan()) over all 65536 wire byte pairs, in both formats, and by
running unit-cbor and unit-bjdata (offline, against the stubbed
test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the MessagePack unsigned-integer writer ladder

The number_integer (non-negative branch) and number_unsigned cases in
write_msgpack() each held their own copy of the fixint/uint8/16/32/64
ladder, kept in lockstep only by a comment ("we used the code from the
value_t::number_unsigned case here"). Both copies mixed union members:
the signed copy compared number_unsigned but wrote number_integer, and
vice versa.

Extract write_msgpack_unsigned(std::uint64_t), mirroring how
write_cbor_head() already avoids the same duplication for CBOR, and
call it from both cases. Each case now reads only its own active
union member. Output bytes are unchanged for the default 64-bit
number types.

#5710 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Unify float marker selection and fix the long double compile error

Four formats picked between a float32 and float64 marker through four
different helper styles: dummy-argument overloads for CBOR and
MessagePack, an std::is_same template for BON8, and a runtime if-chain
on input_format_t for write_compact_float(). With number_float_t set
to long double, to_cbor, to_msgpack and to_ubjson failed inside the
library with "call to 'get_cbor_float_prefix' is ambiguous", while
to_bson kept working because write_bson_double() takes a plain double.

Change write_compact_float() to take the two marker bytes directly
(each of its three callers already knows them at compile time) instead
of an input_format_t it only forwarded, and delete the now-unused
get_cbor_float_prefix(), get_msgpack_float_prefix(),
get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn
the two get_ubjson_float_prefix() overloads into one template. Both
write_compact_float() and get_ubjson_float_prefix() now report an
unsupported number_float_t with a static_assert naming the requirement,
rather than an ambiguous-overload error; the assert lives in the
function body, not the class scope, so to_bson with long double is
unaffected.

Verified with a probe basic_json<..., long double>: to_bson still
compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail
to compile with the new static_assert message.

This changes the text of an existing compile error for users with an
unsupported number_float_t (documented as a public-API-visible change
in #5710).

#5710 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the BJData ndarray writer's dtype dispatch and drop <map>

write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on
every call just to translate the _ArrayType_ name to a dtype marker
(the only reason binary_writer.hpp included <map>), then mapped dtype
to C++ type twice more: once as a switch for the range-check pass and
once as a separate if/else chain for the write pass, with nothing
checking that the two agreed. The caller also ran three at() lookups,
and the callee called value.at(key) about ten more times for the same
three members.

Replace the map with bjdata_ndarray_type_marker(), a plain string
comparison chain (a C++11 constexpr function cannot contain a switch,
so this mirrors binary_reader's own static table style). Replace the
switch/if-chain pair with one write_bjdata_ndarray_elements() that
switches on dtype once and calls a per-type helper -
write_bjdata_ndarray_element<T>() for the eight integer dtypes and
write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag
selecting the range check or the actual write, so the two passes can
no longer disagree on the type. _ArrayType_, _ArraySize_ and
_ArrayData_ are now looked up once into references, and the four
header marker bytes ('[', '$', '#') are written through to_char_type()
like the rest of the UBJSON/BJData writer.

The 'd' (single-precision) rule is left exactly as before, since #5707
is expected to change it separately.

Verified byte-for-byte identical output before/after for every dtype
(including the Draft 2/Draft 3 'byte' fallback and the use_count/
use_type combinations) via a standalone probe, plus round-tripping
through from_bjdata().

Overlaps #5707, which is expected to touch the 'd' dtype case, and
#5518, which is expected to move the write_bjdata_ndarray() call site.

#5710 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Assert that write_bson_document() consumes every calc_bson_sizes() entry

calc_bson_sizes() and write_bson_document() are a hand-synchronized
pair of passes over the same object/array tree, introduced by #5553:
the size pass appends to nested_sizes in visiting order, and the write
pass consumes the table by position with nested_sizes[next_size++].
Nothing checked that the write pass consumed the whole table. If a
future change touched only one of the two passes - for example to skip
or reject an entry - every later size prefix in the document would be
silently wrong.

Add JSON_ASSERT(next_size == nested_sizes.size()) where
write_bson_document() returns, so such a future drift between the two
passes is caught immediately (JSON_ASSERT expands to nothing in
release builds using assert(), and the fuzzers/tests already build
with it enabled). The two passes agree today, so this changes nothing
observable; it only guards against the risk described in #5710 item 5.

Extracting a shared stepper for the two passes (the second half of the
proposed change) is left for a follow-up: it only saves ~30 lines and
the issue asks for it only if the result reads clearly, which needs
more room to get right than a mechanical cleanup pass allows.

#5710 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the BJData lookup tables static functions instead of members

binary_reader held bjd_optimized_type_markers and bjd_types_map as
non-static const members (12 string_t objects for the type-name table),
built and destroyed on every from_cbor/from_msgpack/from_bson/
from_ubjson/from_bon8/from_bjdata call even though only from_bjdata
ever reads them. They also needed the #define/decltype/#undef
workaround from #3637 and two NOLINTNEXTLINE suppressions, and
binary_writer already carries the same two lists in another form
(is_bjdata_excluded_type_marker() and a local std::map in
write_bjdata_ndarray(), the latter removed by the item-4 commit), so
the excluded-marker lists could drift apart.

Replace bjd_optimized_type_markers with static constexpr
is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain
as binary_writer's is_bjdata_excluded_type_marker(). Replace
bjd_types_map with a non-constexpr static bjd_type_name(char_int_type)
switch returning nullptr for an unknown marker (a C++11 constexpr
function cannot contain a switch). Delete both
JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the
NOLINTNEXTLINE suppressions, detail::make_array() (no longer used
anywhere), and the now-unused <algorithm> and <array> includes.

Update the two call sites (the ND-array excluded-type check and the
_ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's
"LUT arrays are sorted" section, which only checked the two tables'
internal ordering, with a check of all 12 type names and all 8
excluded markers against both new functions.

#5711 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read CBOR's 1/2/4/8-byte argument through one helper

parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte
big-endian unsigned integer" ladder four times over:
- twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore
  branch and once, nearly identically, in the ::store branch (~90
  lines to read one integer);
- twice more for container lengths, once for array heads 0x98-0x9B and
  once for map heads 0xB8-0xBB, where the 1/2-byte forms called
  enter_array()/enter_object() directly and the 4/8-byte forms
  additionally went through get_cbor_container_size().

Add get_cbor_argument(std::uint64_t&), reading the width selected by
current & 0x1F via the same get_number() calls as before (so EOF is
reported exactly as before), and route all four sites through it:
- 0xD8-0xDB now read the argument once per branch instead of switching
  on `current` a second time; behavior split cleanly from embedded tags
  0xC0-0xD7 (tag value in the head, no argument to read), which is now
  its own case block that no longer has to fall into the ::store
  switch's "default" case to reach the same tag_pending = true; return
  true; outcome.
- 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always
  going through get_cbor_container_size() (harmless for 1/2-byte
  lengths, which already always fit).

Verified byte-for-byte identical behavior before/after with a
standalone probe covering embedded and multi-byte tags under all three
tag_handler_t settings, a tag over a byte string (subtype path),
truncated tag/length arguments of every width, and array/map lengths
of every width, including the out_of_range.408 "excessive size" case:
same exceptions, same messages, same chars_read, same successful
results.

Left the string/byte-string length ladders in get_cbor_string()/
get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is
expected to touch them separately.

Overlaps #5601 (adds a branch right above the embedded-tag case) and
#5607 (touches the integer cases 0x18-0x1B, which share this ladder's
shape in separate hunks).

#5711 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add leave_container() to match enter_container()

Every container is opened through enter_container(), whose docs
promise that a check placed there runs before every start event. The
close side had no equivalent: the same
"container_stack.pop_back(); dispatch to end_object() or end_array()"
sequence was written out separately in BSON, CBOR, MessagePack,
UBJSON/BJData and BON8, each copying the pattern of keeping an
is_object flag around the pop_back() that would otherwise invalidate
a reference to it. A check needed on close would have had to be added
in five places, and a sixth copy could go unnoticed.

Add leave_container() next to enter_container(), doing the same
pop-then-dispatch, and replace the five sites with it. Each site keeps
its own surrounding logic (BSON's check_bson_document_size() call
before popping, MessagePack's is_object copy used again below,
UBJSON/BJData's remaining-container handling after popping, BON8's
top used again below); only the repeated pop/dispatch line pair is
now shared.

Verified all six binary-format unit suites and unit-regression2's
deep-nesting tests (dependent count/reuse count and the bjdata ndarray
depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan.

Overlaps #5601, which is expected to add a sixth close site in its own
skip loop; that site can route through leave_container() too once it
lands.

#5711 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop passing the input format to sax_parse() when the reader already has it

binary_reader's constructor stores the format in the input_format
member, and sax_parse(format, sax_, strict, tag_handler) took the same
value again purely to dispatch on it. Every in-tree caller passed the
same value both times (all 16 from_cbor/from_msgpack/from_ubjson/
from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three
public basic_json::sax_parse() overloads), so nothing was broken
today, but a caller of the detail class directly (only reachable via
JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could
pass a mismatched pair - say bjdata to the constructor and ubjson to
sax_parse - and dispatch on one format while applying the other
format's rules; the default-constructed input_format_t::json reader
would additionally hit JSON_ASSERT(false) in exception_message() on
its first error.

Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the
existing overload with the stored input_format, and switch every
caller to it: the 16 from_*() sites (keeping their
`// cppcheck-suppress[accessMoved]` comments) and the three
basic_json::sax_parse() overloads, all of which already had the format
available from their own `format` parameter. The four-argument overload
is kept for anyone still calling it, now with
JSON_ASSERT(format == input_format) so a mismatch fails immediately
in a debug build (assert-enabled binaries, including the fuzzers and
test suite) instead of misbehaving; verified with a probe that
constructs a reader for one format and calls the explicit overload
with another, which aborts on that assertion as expected.

Removing or asserting against the constructor's input_format_t::json
default, which would affect direct detail users, is left as a separate
decision per #5711 item 5.

Overlaps #5601, which is expected to add an AllowRecovery template
parameter to sax_parse() and touch these same call sites in json.hpp.

#5711 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks

get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently
sized signed integer and then repeated the same "reject negative with
error 113" check; only 'L' additionally checked value_in_range_of for
the out_of_range.408 case. Any change to that error path had to be
made four times.

Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read,
the negative check and the range check once, and route all four
markers through it. The range check is a no-op for 'i'/'I'/'l' (their
values always fit std::size_t) and only live for 'L' on a 32-bit
std::size_t target, matching today's behavior exactly.

In the ndarray dimension-product loop, the preceding loop already
returns early on any zero dimension and result starts at 1, so `i > 0`
in the pre-multiplication overflow check was always true, and
`result == 0` in the post-multiplication check could not be reached
either: two positive factors whose product does not overflow (as the
pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&`
and narrow the post-check to `result == npos`, the one case the
pre-check cannot rule out (an exact, non-overflowing match with the
sentinel reserved for unknown-size containers), with a comment
explaining why.

Verified byte-for-byte identical behavior before/after with a
standalone probe covering negative counts for every marker, a matching
positive count, and ndarray inputs, plus the full unit-ubjson and
unit-bjdata suites (same assertion counts as before this change).

Overlaps #5601 (rewrites the four parse_error calls and the overflow
checks touched here) and #5607/#5707 (touch neighboring lines in the
same functions).

#5711 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop redundant format parameter and dummy float argument (review)

binary_reader::sax_parse(format, ...) only ever had to equal the format
given to the constructor, which it asserted. With every caller already
on the format-less overload, remove the four-argument overload and
dispatch on the stored input_format directly. binary_reader is a
detail class, so this is not a public API change.

get_ubjson_float_prefix() took a value only to deduce its type; make
the type an explicit template argument instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:48:32 +02:00
Niels Lohmann
7d7055ec50 Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations

Constructing a basic_json from another specialization (json to
ordered_json or back, also via get<ordered_json>()) converted every
container with its range constructor, which calls the converting
constructor for each element. The call stack therefore grew with every
nesting level, and a value nested some 30,000 levels deep overflowed it.

The conversion now bounds its descent the way the copy constructor does
since #5387: the first 128 levels are converted exactly as before, and
below that convert_iteratively() finishes the value with an explicit
stack. It builds each container bottom-up from its converted elements
with the container's range constructor, so member order and keys that
become equal are handled as before, and it gives a value its type only
once its container exists, so an exception leaves nothing behind that
cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions
(JSON_DIAGNOSTIC_POSITIONS) are set for every value.

Converting a null value no longer resets its positions: the constructor
assigned null to a value that already was null, which swapped in the
positions of the temporary.

Fixes #5650.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Explain why converting null keeps positions and why next is a reference

Review feedback on #5723 (gregmarr): clarify in comments that the
converting constructor has already copied the positions of val, which
the null case keeps like every other case, and that next must be a
reference into pending so that ++next advances the stored iterator.

Comments only; no code change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Refer to recursion_depth_limit() in the convert_structured() docs

The comment still named nesting_depth_limit, which #5637 removed on
develop in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Advance the pending iterator through pending.back() and shorten the null comment

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:36:04 +02:00
Niels Lohmann
3b6ae43c53 Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions

basic_json can be constructed from std::tuple<json&>, which it turns into
a one-element array. Because of this, std::tuple picks its converting
constructor that converts the whole source tuple instead of the
element-wise one. As a result, std::tuple<const json&> built from
std::forward_as_tuple(j) binds to a temporary (a compile error with libc++,
a dangling reference with other standard libraries), and std::tuple<json>
built the same way holds [j] instead of a copy of j.

The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option
JSON_DisableTupleReferenceConversion) removes the conversion from a
one-element tuple holding a reference to the same basic_json type, so
std::tuple converts element-wise. It is off by default, so existing
behavior is unchanged. It does not change any function body and therefore
is not part of the ABI tag.

Fixes #2226

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert one-element tuples to arrays on every compiler

to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }.
With a single element that is itself a basic_json, Apple clang 15 and 16
treat j = {x} as a copy of x, so std::tuple<json>{true} became true
instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new
checks in unit-disable-tuple-reference-conversion and unit-regression2.

The one-element overload that already handles
JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a
[string, value] element) explicitly, the same way the initializer-list
constructor does. Use it unconditionally. The output is unchanged on
compilers that already wrapped the element.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip json reference tuple tests on clang < 4 and GCC < 5

ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could
not compile the new tuple tests. Creating a std::tuple of basic_json
references, e.g. std::forward_as_tuple(j), makes these compilers
instantiate basic_json's conversion operator for libstdc++'s internal
tuple bases, which fails hard. This happens with and without
JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these
compilers, not of the new option.

Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail,
clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the
checks that create such tuples; the is_constructible checks still run.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:53:56 +02:00
Niels Lohmann
6a073dbae4 Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().

Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.

Fixes #5652.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:34 +02:00
Niels Lohmann
fdcc569eee Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.

Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.

Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).

Fixes #5666.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:31 +02:00
Niels Lohmann
6fc0d501f3 Move the user-defined string literals to <nlohmann/json_literals.hpp> and add JSON_NO_AUTOMATIC_UDLS (#5610)
* Add JSON_NO_UDLS to leave out the user-defined string literals

The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mention JSON_NO_UDLS in the list of exported module symbols

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the user-defined string literals to <nlohmann/json_literals.hpp>

Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.

The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test

- Suppress clang-tidy misc-header-include-cycle on the intentional mutual
  include of json.hpp and json_literals.hpp.
- Use operator"" _json with a space for GCC 4.8 in the test's detection
  aliases, as the header does.
- Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Declare the literal operators through a local macro

The GCC 4.8 spacing condition was repeated for both operator definitions
and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix)
now selects operator""##suffix or operator"" suffix in one place and is
undefined at the end of json_literals.hpp.

Suggested by gregmarr in review.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:17 +02:00
Niels Lohmann
67435c9c7e Give a deep copy its type only after its container exists (#5721)
When copying a value nested deeper than 128 levels, and an allocation
fails while an inner array or object is being copied, the partially
built copy ended up with an element typed array/object but holding a
null pointer. That element was already a fully constructed member of
its parent's container, so destroying the parent during stack
unwinding dereferenced the null pointer (release builds) or failed
assert_invariant() (debug builds), instead of letting std::bad_alloc
reach the caller.

copy_iteratively() set a pending worklist element's type right after
popping it, before the next loop iteration created its container in
copy_array_level()/copy_object_level(). Move that type assignment into
those two functions, right after the container is successfully
created, and drop the premature one in copy_iteratively(), so a
half-built element stays a null value - as copy_shallow()'s comment
already promised - until it can safely hold one.

Add a regression test to tests/src/unit-allocator.cpp that copies a
value nested 130 levels deep (both arrays and objects, with a
std::map- and an ordered_map-backed object_t) and fails every
allocation of the copy in turn: each attempt must throw std::bad_alloc
without crashing, and the source must stay unchanged.

Fixes #5640.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:00 +02:00
Niels Lohmann
2ea6d8c127 Require the found key to equal the looked-up key when comparing objects (#5720)
For an object type whose comparator treats unequal keys as equivalent
(for example a std::map with a case-insensitive comparator),
compare_iteratively() looked up a mismatched left key in the right
object with find(), which uses the object's own comparator, and
accepted whatever entry it found without checking that the keys are
actually equal. A case-insensitive comparator then found "KEY" for
"key", so two objects nested past the recursion bound (or at every
depth with JSON_NO_THREAD_LOCAL) could compare equal even though the
object type's own operator== - and basic_json itself, below the bound
- consider them different.

Accept the found entry only if its key equals (not just compares
equivalent to) the looked-up key.

Fixes #5655.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:56 +02:00
Niels Lohmann
66877675b1 Check the iterator range for binary values in basic_json(first, last) (#5719)
basic_json(first, last) treated value_t::binary like the structured
types (array, object) in the range check, so it always copied the
whole binary value regardless of the iterators, even for an empty
range such as (b.end(), b.end()). The other primitive types (number,
boolean, string) already reject such a range with
invalid_iterator.204, and erase(first, last) already does the same
for binary values, so this made the constructor inconsistent with
both. Move case value_t::binary into the group of checked primitive
types.

Also update the two matching passages in basic_json.md that describe
overload 7, and add a version-history note.

Fixes #5670.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:52 +02:00
Niels Lohmann
7fd6895788 Hide a discarded container's content from the parser callback (#5706)
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.

The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.

Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.

Fixes #5643.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:48 +02:00
Niels Lohmann
6ae17630a4 Reject integral keys for contains(), find(), and count() at compile time (#5705)
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is
a null pointer constant, so it converts to a null const char*, and the
overloads taking const typename object_t::key_type& accepted it by
constructing a std::string from that null pointer, which is undefined
behavior (a crash with both libc++ and libstdc++). value(0, default_value)
had the same problem in C++11, where the object comparator is not
transparent.

Add deleted overloads for integral arguments to contains(), find()
(const and non-const), count(), and value() so that these calls are
compile errors in every supported language mode instead of crashing.
Calls with string, string_view, json_pointer, and size-typed element
access (at(), operator[](), erase()) are unaffected.

Fixes #5657.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:45 +02:00
Niels Lohmann
5bd766aa50 Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.

calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.

Fixes #5675.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:41 +02:00
Niels Lohmann
4bb1b14b06 Trim the compiler-appended NUL from wide/UTF string literals too (#5702)
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.

Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.

Fixes #5658.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:37 +02:00
Niels Lohmann
bfe0f32d71 Fix std::terminate and null pointer access in input_stream_adapter (#5699)
Parsing from a std::istream crashed in two unusual but valid stream
states, both in input_stream_adapter:

- With eofbit in the stream's exceptions() mask, get_character() sets
  eofbit via is->clear(), which throws std::ios_base::failure. While
  that exception unwinds, ~input_stream_adapter() called clear() again
  to reset eofbit, which is still set and still in the exception mask,
  so it throws a second time out of the (implicitly noexcept)
  destructor and std::terminate() is called. The destructor now only
  calls clear() if a bit other than eofbit remains set, so the first
  exception can propagate normally.
- For an std::istream without a stream buffer (rdbuf() == nullptr,
  e.g. std::istream(nullptr)), the constructor stored the null
  pointer without checking it, and get_character() dereferenced it.
  input_adapter(std::istream&) now throws parse_error.101 for such a
  stream, the same as it already does for a null FILE* or char*.

Added regression tests to unit-deserialization.cpp and, for the
JSON_PRECISE_STREAM_POSITION variant of get_character(), to
unit-precise-stream-position.cpp; both crashed before this fix.
Documented the two exceptions in parse.md and operator_gtgt.md.

Fixes #5646.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-30 20:07:33 +02:00
Niels Lohmann
44a88d85be Fix NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json message (#5698)
from_json built its out_of_range.410 message with "..." + j.dump(). If the
unmatched value is (or contains) a string with invalid UTF-8, that dump()
itself throws type_error.316, so the caller got type_error.316 instead of
the documented out_of_range.410; such strings can reach get<Enum>()
unvalidated, e.g. from from_cbor()/from_msgpack(). With a custom string_t,
j.dump() returns that type, and "const char*" + string_t does not compile
unless the type happens to provide operator+, so the macro failed to
compile for such types.

Build the message with detail::concat(), which appends any type exposing
data()/size() and always yields a std::string, and dump with
error_handler_t::replace so building the message itself cannot throw.

Added regression tests: an invalid-UTF-8 case in the existing strict-enum
test in unit-conversions.cpp, and a strict-enum use with alt_string (the
custom string_t from unit-alt-string.cpp) to cover the compile failure.

Fixes #5667.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:29 +02:00
Niels Lohmann
5b11a0282c Exchange the CustomBaseClass subobject in basic_json::swap() (#5697)
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap
overload that forward to it) only exchanged m_data.m_type/m_data.m_value,
leaving each value's json_base_class_t subobject in place. This is
inconsistent with the copy and move constructors and copy assignment,
which all carry the base class along with the value, so after
a.swap(b) any metadata stored in a CustomBaseClass ended up attached to
the wrong value. Algorithms that mix swap() with moves, such as
std::sort, scrambled the metadata across the whole container.

Fix the member swap() to also exchange the json_base_class_t subobject
and extend the noexcept specifications of swap() and the friend swap()
accordingly.

Fixes #5653.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:24 +02:00
Niels Lohmann
bfea6f36d3 Fix to_msgpack() reading the inactive number union member (#5694)
* Fix to_msgpack() reading the inactive number union member

basic_json stores number_integer and number_unsigned in a union, and
number_unsigned_t only has to be at least as wide as number_integer_t
(with the default types, both are 64-bit and have the same
representation). When number_integer_t is narrower, write_msgpack()
read the wrong union member in two places:

- The number_unsigned case wrote number_integer's bits instead of
  number_unsigned's, silently writing the wrong value whenever it
  did not fit in number_integer_t.
- The number_integer case (non-negative branch) picked the encoded
  width by comparing number_unsigned's bits, which is undefined
  behavior, though the value written was still number_integer's, so
  at worst a too-wide encoding was chosen.

Read the active member in both cases, like the other binary writers
(CBOR, UBJSON, BJData, BSON, BON8) already do.

Fixes #5644.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cast number_integer to number_unsigned_t only once in to_msgpack()

Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:19 +02:00
Niels Lohmann
e444a66276 Copy values before inserting an initializer list into an array (#5693)
* Copy values before inserting an initializer list into an array

insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).

Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.

Fixes #5656.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the reserve_array helper in the initializer_list insert fix

The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.

Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:16 +02:00
Niels Lohmann
1151826508 Take diff()'s fast path unless the object type reorders members (#5691)
* Take diff()'s fast path unless the object type reorders members

For every object type except an insertion-ordered one like ordered_map,
diff() no longer produced a member-by-member patch when target had a key
that sorts before a key the two objects share: it fell through to the
slow path, which removes every member of source and re-adds every member
of target, instead of just adding the new key.

#5465 added an order check to require the fast path to also reproduce
target's member order, needed because ordered_map's patch()-driven "add"
appends a new member at the end. The check compared the common keys'
order between source and target and also required that every added key
come after every common key in target's order ("new_keys_form_suffix").
The comment above it argued this check is always true for std::map, and
that reasoning is correct for the order of the common keys themselves,
but not for new_keys_form_suffix: a std::map iterates in sorted key
order, so a new key that sorts before an existing common key is
enumerated between common keys, making new_keys_form_suffix false even
though std::map's own key order does not need reordering at all - it
places every member itself, regardless of insertion history, so a
member-by-member diff already reproduces target's iteration order.

Only require the order check for an object type that keeps insertion
order, using the same detail::is_ordered_map trait the library already
uses to recognize such an object type in set_parent(). Every other
object type - std::map in key order, a hash map in an order its
operator== ignores - always takes the fast path.

Added a regression test to unit-json_patch.cpp: the issue's example now
yields a single "add" op for json, while ordered_json still takes the
slow path to reproduce target's member order.

Fixes #5639.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Only track the target key order in diff() for insertion-ordered objects

common_keys_target_order and new_keys_form_suffix are only read when object_t keeps its members in insertion order; skip building them otherwise. Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:10 +02:00
Niels Lohmann
6218aa212b Throw std::length_error for operator[](SIZE_MAX) instead of corrupting the array (#5687)
For idx == SIZE_MAX, the non-const array operator[] computed the new size as
idx + 1, which wraps to 0. resize(0) then emptied the array, and the
subsequent operator[](idx) on the now-empty vector wrote one element before
its buffer. Every other too-large index (e.g. SIZE_MAX - 1) already went
through resize(), which throws std::length_error and leaves the array
unchanged; SIZE_MAX was the one value for which the overflow bypassed that
safety net.

Add a guard that throws std::length_error before computing idx + 1 when idx
is the largest representable size_type value, so the array is left
unchanged, matching the exception vector::resize() already throws for
smaller (but still too large) indices.

Fixes #5647.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:06 +02:00
Niels Lohmann
42f89e3130 Let to_json(std::optional<T>) propagate exceptions from T's to_json (#5684)
* Let to_json(std::optional<T>) propagate exceptions from T's to_json

The overload was marked noexcept even though its body assigns *opt to
the JSON value, which calls T's to_json (or allocates for std::string,
std::vector, or json). Any exception from there -- a user-defined
to_json reporting an error, or std::bad_alloc -- called std::terminate()
instead of propagating. The noexcept also made basic_json's converting
constructor noexcept(true) for std::optional<T>, so json j = opt; could
not report the error either.

Fixes #5642.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: make to_json(std::optional<T>) conditionally noexcept

GCC's -Wnoexcept (an error in ci_test_gcc) fired at to_json_fn's
noexcept(noexcept(to_json(j, val))): after dropping the unconditional
noexcept, to_json(std::optional<int>) had no exception specification
although GCC could prove its body cannot throw. It also made
json(std::optional<int>) lose its noexcept.

Declare the overload noexcept exactly when assigning the contained
value to the JSON value is (std::is_nothrow_assignable<BasicJsonType&,
const T&>), which is what the body does. std::optional<int> is noexcept
again; a T whose to_json may throw still propagates the exception.
Static assertions in the test check both cases.

The test's throwing to_json triggered -Wmissing-prototypes and
-Wmissing-noreturn (clang) and -Wmissing-declarations and
-Wsuggest-attribute=noreturn (GCC). Move the type and its to_json into
an anonymous namespace, mark the function [[noreturn]], and compile
them only without JSON_NOEXCEPTION, like the test that uses them.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:02 +02:00
Niels Lohmann
6d7845d207 Fix deprecated json_pointer/string operator== warning in value() (#5683)
value(KeyType&&, default) is constrained on is_comparable_with_object_key,
which passes KeyType as a reference. is_comparable's dispatch on
is_json_pointer_of<A, B> only matches a json_pointer as a plain type or a
plain reference, so a const-qualified reference (as produced when KeyType
is deduced from a json_pointer argument) fell through to
is_comparable_no_json_pointer, which instantiates the deprecated
json_pointer/string comparison operators. This made ordered_json's
transparent comparator (and any transparent comparator on a custom string
type) warn under -Wdeprecated-declarations when calling
value(json_pointer, default), even though no such comparison is ever
performed. at() was already fixed for this in #5289, which does not use
is_comparable_with_object_key.

Strip references and cv-qualifiers with uncvref_t before the
is_json_pointer_of dispatch, so any reference-to-json_pointer is
recognized regardless of qualifiers.

Fixes #5664.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:59 +02:00
Niels Lohmann
678fd3017b Fix element path for map/unordered_map JSON_DIAGNOSTICS errors (#5681)
When converting a JSON array to std::map or std::unordered_map with a
non-string key, each element must itself be a [key, value] array. If an
element is not an array, from_json() threw type_error 302 with the outer
array's value (&j) as the exception context, so with JSON_DIAGNOSTICS
enabled the message pointed at the whole array instead of the offending
element (e.g. "(/outer/m)" instead of "(/outer/m/2)"), even though the
message text already described the element's type.

Both from_json() overloads now pass the element (&p) as the context, so
the reported JSON Pointer matches the type named in the message, the
same way std::vector<std::vector<T>> and similar conversions already do.

Fixes #5668.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:55 +02:00
Niels Lohmann
fc4c9c3446 Fix clear() to also reset the subtype of a binary value (#5680)
* Fix clear() to also reset the subtype of a binary value

clear() on a binary value cleared the bytes but left the subtype
untouched, so the result was not equal to a default-constructed
binary value even though the documentation says clear() has the same
effect as *this = basic_json(type()). The fix calls
byte_container_with_subtype::clear_subtype() alongside the existing
clear() call.

Extended the "filled binary" clear() test in unit-modifiers.cpp with
a case that uses a subtype, since the existing cases only covered
binary values without one.

Fixes #5669.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the table alignment in clear.md

Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:51 +02:00
Niels Lohmann
68beba727c Fix from_json() for enums with underlying type bool (#5679)
get_arithmetic_value() rejects boolean_t, so the default from_json()
for enums failed to compile for an enum whose underlying type is
bool (e.g. enum class Flag : bool { off, on }), even though the
matching to_json() serializes such enums as an unsigned number.

Read the underlying value through number_unsigned_t in that case,
matching what to_json() writes, then cast back to the underlying
type before constructing the enum.

Fixes #5671.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:46 +02:00
Niels Lohmann
6a8a7735ed Do not throw in contains() for an empty array reference token (#5614)
* Do not throw in contains() for an empty array reference token

json_pointer::contains() rejected malformed array indices, but an empty
reference token (e.g. "/a/" where "a" is an array, or "/" on an array)
passed every check and reached array_index(), which throws
out_of_range.404. contains() must not throw (cf. #5395), so it now
returns false for an empty token. at() still throws out_of_range.404.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: bind j_nested_const by reference in the json_pointer test

clang-tidy (ci_clang_tidy) flagged the new test with
performance-unnecessary-copy-initialization: the local copy
j_nested_const of j_nested is never modified. Bind it as a const
reference instead; it still exercises the const overloads of at() and
contains().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:39 +02:00
Niels Lohmann
65af260117 Move instead of deep-copy ordered_json values when an object grows (#5609)
* Move instead of deep-copy ordered_json values when an object grows

ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>.
With a std::string key, that pair is not nothrow move constructible (the
const key has to be copied), so std::vector copies every element when it
reallocates. For ordered_json, this deep-copies every member value an
object already holds, including whole nested subtrees, on each growth
step.

Grow the storage in ordered_map instead, copying the keys and moving the
values. This happens in two phases, so the strong exception guarantee is
kept without try/catch. The first phase may throw, but only touches a
temporary buffer: it copies the keys, value-initializes the values, and
constructs the new element. The second phase moves the values (noexcept)
and swaps the buffers. Because the new element is constructed before any
value is moved, arguments that refer to elements of the container stay
valid, as with std::vector. Types that cannot take this path keep the
std::vector behavior.

Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3):
twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 ->
177 ms, canada unchanged. The number of allocations for twitter and
citm_catalog drops by two thirds.

Also add ParseStringOrdered rows to the benchmarks, and document the
growth behavior and the exception safety of ordered_map.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: skip std::pair noexcept assumptions on EDG-based compilers

ci_icpc and ci_nvhpc failed to compile unit-ordered_map.cpp: the static
assertion that std::pair<const std::string, ordered_json> is not nothrow
move-constructible fails there. The EDG front end (Intel icpc 2021.10,
NVIDIA nvc++ 25.5) considers the defaulted move constructor of
std::pair<const Key, T> noexcept even if copying Key can throw. With these
compilers, std::vector already moves such elements itself when it grows,
and ordered_map correctly leaves growing to it.

The same misjudgement makes std::vector call std::terminate when a key copy
throws during growth, so the exception-safety test with throwing_key would
abort on these compilers as well.

Skip the static assertion and the exception-safety section when __EDG__ is
defined. Verified with icpc 2021.10 (-std=gnu++11) and nvc++ 25.5 (C++11 and
C++17) on Compiler Explorer: unit-ordered_map and unit-disabled_exceptions
build and pass.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:34 +02:00
Niels Lohmann
d268eaa693 Convert floats with Eisel-Lemire when std::from_chars is unavailable (#5617)
* Convert floats with Eisel-Lemire when std::from_chars is unavailable

Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).

The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.

json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: unused parse result and find() == npos in the Eisel-Lemire tests

GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...))
because parse() is [[nodiscard]]; assign the result to a dummy json as the
other tests do. clang-tidy flagged longer.find('.') == npos with
abseil-string-find-str-contains; store the position in a variable first.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:26 +02:00
Niels Lohmann
633de8e44b Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:

- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
  calls whose result was discarded. Check the result, like the other
  resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
  callbacks, which cannot throw but were not declared noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:20:43 +02:00
Niels Lohmann
fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann
9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00
Niels Lohmann
373005f7ac Fix MSVC: avoid reserving by faked size in MessagePack size tests (#5604)
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).

Write into a caller-owned vector instead, so nothing is reserved.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:57:16 +02:00
Niels Lohmann
509c07041f Fix CI: clang-tidy and clang/libstdc++ 10 in the MessagePack size tests (#5599)
The tests added by #5515 fail two ways on develop:

- clang-tidy reports the size() overrides of huge_string and huge_binary
  (readability-convert-member-functions-to-static) and the non-const
  test value (misc-const-correctness); mark them like the #5584 types
- clang with libstdc++ 10 cannot compile the file for C++17: the
  std::filesystem::path conversion considered for huge_string, a class
  derived from std::string, is ambiguous. Guard it with
  JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same
  reason, and define that macro before both test blocks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 17:50:19 +02:00
Niels Lohmann
1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00
Niels Lohmann
f682cd2ef1 Skip the #5515 MessagePack size tests on 32-bit platforms (#5590)
The tests fake a container size of UINT32_MAX + 1, which does not fit
into a 32-bit std::size_t: MSVC rejects the truncation (C4305/C4309
with /WX), and clang-cl wraps the size to 0 so nothing throws. Guard
them with SIZE_MAX > UINT32_MAX like the tests from #5584.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 15:58:59 +02:00
Kartikey Negi
6bd106893a Fix CBOR tag handling in cbor_tag_handler_t::store for non-binary items (#5559)
When using cbor_tag_handler_t::store, tags 0xD8-0xDB previously assumed
that the tagged item was a byte string, unconditionally attempting to
parse binary data and failing on valid CBOR documents containing tags
applied to integers, strings, arrays, or objects (such as self-describe
tag 55799).

Check whether the tagged data item is a byte string (0x40-0x5B or 0x5F).
If it is a byte string, store the subtype on the binary value as before.
Otherwise, iteratively process the tagged value in the driver loop using
item_read so that chained tags do not consume native stack space.

Part of #5316.

Signed-off-by: ReturnKartikey <kartikeynegi2000.work@gmail.com>
2026-09-27 14:28:55 +02:00
bucketbase26
98e00d22e5 Cut test suite runtime in binary roundtrips and integer sweeps (#5519)
* Cut test suite runtime in binary roundtrips and integer sweeps

The Linux CI jobs pass --no-skip, so skip() does not help there.
Parse each corpus file once in the binary roundtrip loops instead of
four times. Sample the 16-bit integer ranges with stride 7 (still hits
every low byte) and always keep the endpoints.

Also drop the 5M-node parse test to 500k, which still covers the
non-recursive destructor, and move jeopardy.json into its own skipped
test so the cheaper binary-format size checks actually run.

See #5418.

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>

* Drop useless int32_t casts in the sampled integer loops

ci_test_gcc compiles with -Werror=useless-cast. On that compiler
int32_t is int, so static_cast<int32_t> of the loop bound is an
error. The bounds are already int, and the sampled values do not
change.

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>

* Revert unit-binary_formats.cpp to develop and fix comment

Revert tests/src/unit-binary_formats.cpp to its develop state.
The test-case split made valgrind jobs slower instead of faster,
because the cheaper corpus files (canada/twitter/citm/sample)
now ran under valgrind where they never did before.

Fix the next_integer_sample comment: the function has no 'first'
parameter, so describe what the function actually does.

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>

---------

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
2026-09-27 14:28:38 +02:00
Niels Lohmann
f7972970a4 Throw instead of writing MessagePack lengths beyond UINT32_MAX (#5584)
* Throw instead of writing MessagePack lengths beyond UINT32_MAX

MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.

The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the CI failures of the MessagePack length check

- mark to_msgpack_length's value as used when exceptions are disabled
  (-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
  documentation style check requires
- create the test's string value from its type: constructing it from a
  beyond_uint32_string_t considers the std::filesystem::path conversion,
  which libstdc++ 10 reports as ambiguous for a class derived from
  std::string (clang 13)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip the MessagePack string length test for clang with libstdc++ 10

C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the MessagePack string test type and its alias in one block

astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 14:28:16 +02:00
Niels Lohmann
6178982b8d Compare unordered objects by key below the nesting bound (#5582)
* Compare unordered objects by key below the nesting bound

Values nested deeper than the nesting bound are compared without the
call stack, walking both objects entry by entry. Two equal objects of a
type that enumerates its entries in no fixed order - std::unordered_map,
say - can be walked in different orders, so they compared unequal, and
a deep copy compared unequal to its original. std::unordered_map's own
operator== does not depend on the order, which is what applies above the
bound.

Where the keys differ, equality now finds the entry by its key instead.
An ordering, and ordered_map, whose operator== compares its entries in
sequence, still decide by the key.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Test unordered object equality without std::unordered_map

basic_json<std::unordered_map> instantiates std::pair<const string,
basic_json> while basic_json is still incomplete. The standard does not
require std::unordered_map to support that, and libstdc++ 6 to 9 as well
as the EDG front ends of icpc and nvc++ reject it, which broke the build
of unit-comparison on those CI jobs.

The test now uses an object type derived from std::map (which, as the
default object type, works everywhere) whose comparator orders keys
ascending or descending as chosen at construction, and whose operator==
does not depend on the order of the entries - the property of
std::unordered_map the test is about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Compare the test object type's entries with std::all_of

clang-tidy (readability-use-anyofallof) asked for std::all_of instead of
the loop in unordered_object_t's operator==. The entry type is spelled
out, as C++11 needs typename for base_type::value_type and C++20
reports it as redundant.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 14:21:58 +02:00
Niels Lohmann
85f8b21e1c Add tests for uncovered code paths (#5581)
* Add tests for uncovered code paths

Cover code the test suite did not reach, found from the Coveralls report
of develop and a local coverage run of HEAD:

- dump() of every kind of value below the bound of the recursive descent
  (pretty-printed objects, binary values, discarded values, scalars), and
  flushes of the escape and write buffers mid-string and mid-binary
- the iterative comparison: objects with different keys, containers that
  are a prefix of each other, and elements that cannot be ordered, each
  both at the top level and below the nesting bound
- SAX handlers that stop at any event, including the end of a nested
  container, in the BSON, CBOR, MessagePack, UBJSON and BJData readers
- from_bson/cbor/msgpack/ubjson/bjdata returning a discarded value
  through the iterator and pointer overloads
- JSON Patch, diff, merge_patch and update(..., true) on ordered_json
- smaller gaps: get_allocator(), to_ubjson/to_bjdata into a string,
  value() with an unresolvable JSON pointer, integer/float comparison
  below the integer range and with negative fractions, conversion to a
  custom binary type, std::formatter::parse on a spec without '}',
  unescape() of a lone '~', and the callback parser's start_array()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cover more paths that were thought unreachable

- parse_float_fast() declining malformed or inexact input, called
  directly since the lexer only passes well-formed numbers to it
- a UTF-16 high surrogate followed by a unit above the low surrogates
- self-assignment of a const_iterator
- a truncated CBOR string read through non-contiguous iterators
- serializing a long double under the de_DE locale, which undoes the
  locale's decimal point and thousands separator
- values read from a binary format carrying no diagnostic positions,
  with and without a parser callback

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the CI failures of the new coverage tests

- declare the self-assignment reference const (misc-const-correctness)
- expect the (/path) prefix that JSON_DIAGNOSTICS adds to the messages
  of the failing ordered_json patch operations

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Expect the byte range JSON_DIAGNOSTIC_POSITIONS adds to the patch errors

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Build the expected dump of the nested-object test with +=

clang-tidy (performance-inefficient-string-concatenation) reported the
chain of operator+ calls that assembled the expected indented output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Compare the BJData and UBJSON test outputs byte by byte

Building a std::string from the byte vector converts each byte
implicitly, which -fsanitize=integer reports for bytes of 0x80 and
above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 14:19:22 +02:00
Niels Lohmann
4fa95d9810 Remove unreachable branches from the binary writer (#5583)
Coverage reported conditions in the binary writer that can never be
false, and marked the code behind them with LCOV_EXCL. Remove them
instead of excluding them:

- CBOR writes the length of a string, binary value, array, or object
  exactly like an unsigned integer, only with another major type. One
  function, write_cbor_head(), now writes both, so the integer tests
  cover every width and the four excluded 64-bit length branches are
  gone.
- A last `else if` whose condition holds for every remaining value
  (an unsigned value at most UINT64_MAX, a signed one in the range of
  int64_t) is now a plain `else`.
- Whether a signed integer fits into an int64 for UBJSON and BJData is
  decided by its type at compile time. Only an integer type wider than
  64 bits gets a range check and the high-precision fallback.
- The private get_impl(boolean_t*) was never called.

The UBJSON type prefix 'H' of an optimized container of unsigned
integers beyond the range of int64 was reachable although excluded; it
is tested now.

The output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 14:18:46 +02:00
Dadi Reddy Sai Praneeth Reddy
fe4a544c7e handled when size exceed uint32 (#5515)
* handled when size exceed uint32

Signed-off-by: dsp0redy <saipraneethreddy.dadireddy@gmail.com>

* addressed review comments

Signed-off-by: dsp0redy <saipraneethreddy.dadireddy@gmail.com>

* updated unit test

Signed-off-by: dsp0redy <saipraneethreddy.dadireddy@gmail.com>

* added amalgamation patch

Signed-off-by: dsp0redy <saipraneethreddy.dadireddy@gmail.com>

---------

Signed-off-by: dsp0redy <saipraneethreddy.dadireddy@gmail.com>
2026-09-27 14:17:32 +02:00
Niels Lohmann
95e9a5931c Write BSON in linear time, without recursing per nesting level (#5553)
* Write BSON in linear time, without recursing per nesting level

to_bson() had two problems with nested values:

- It recursed once per nesting level, so a value nested deeply enough -
  100,000 levels on an 8 MiB stack - exhausted the call stack and
  terminated the process, although parse() accepts such values without
  complaint.
- BSON prefixes every document and array with its length. The writer
  computed that length by walking the entire value below it, again for
  every nested document it wrote, which made serializing O(size x depth).
  A 200-level document took 30 ms instead of 1.

Both passes are now iterative, and each length is computed exactly once:

- calc_bson_sizes() computes the length of every document and array in
  one pass, each from the lengths of its entries, into a table ordered
  the way they are written.
- write_bson_document() then writes the document, taking each length from
  the table.

Everything observable is unchanged, as a differential test against
develop confirms byte for byte:

- The same bytes are written.
- A key containing U+0000 still throws out_of_range.409 for the same
  first key, with the same diagnostics path, before anything is written.
- A document too large for BSON still throws out_of_range.412 before
  anything is written.
- A binary subtype above 255 still throws out_of_range.415 after the
  same partial output.

Only the enclosing objects and arrays are kept on a stack, so a flat
document allocates nothing for it. Measured against develop (clang -O3,
median of 201 runs): flat objects unchanged, flat arrays 37% faster (the
array length was computed twice), a nested 3,000-object document 2x
faster, a 200-level document 33x faster.

to_bson.md documented the quadratic complexity since #5334; it is linear
again.

Fixes #5392 for BSON, and #5308.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not require a default-constructible string_t in the BSON writer

GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default
constructor; develop never default-constructed string_t here either.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Let the BSON index-name helper only fill its output parameter

It returned a reference to the string it filled, so callers held a second
name for index_name. Addresses review feedback.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 21:56:39 +02:00
Niels Lohmann
1e44262091 Make JSON_STRICT_NUL_HANDLING part of the ABI tag (#5560)
* Make JSON_STRICT_NUL_HANDLING part of the ABI tag

JSON_STRICT_NUL_HANDLING (#5534) changes the bodies of inline functions:
the lexer's handling of '\0' and input_adapter() for char arrays. So
translation units compiled with and without it define the same functions
differently, an ODR violation - the case the ABI tag exists for, as with
JSON_BRACE_INIT_COPY_SEMANTICS (_bics). It now appends _snul to the inline
namespace. The macro is new in 3.13.0, so no existing namespace changes.

Its default moves to abi_macros.hpp, and it is only #undef'd without
JSON_TEST_KEEP_MACROS, as for the other ABI macros. The ABI config tests,
the namespace docs, the macro's docs and the Natvis file cover the new tag.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 21:54:11 +02:00
Niels Lohmann
c60a0bc336 Allocate the deep copy's key scratch space with the provided allocator (#5573)
* Allocate the deep copy's key scratch space with the provided allocator

The iterative deep copy builds each object's keys in a temporary vector of
key/value pairs before handing them to the object's range constructor. That
vector holds basic_json values, so like the values themselves it now uses
AllocatorType instead of std::allocator.

Also document that AllocatorType covers the JSON values, while most
temporary storage still uses std::allocator.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Count allocate_at_least in the scratch-counting test allocator

From C++23 on, libc++'s containers allocate through allocate_at_least when
the allocator has one. The test allocator inherited it from std::allocator,
so the scratch allocations were not counted and the test failed on Xcode.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 21:53:49 +02:00
Niels Lohmann
d19f7f5dce Fix BSON conformance issue (#5185)
* 🐛 fix BSON conformance issue

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 🐛 fix BSON conformance issue

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)

from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).

Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.

There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).

Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* ⚡ validate only newly read bytes of binary-format strings

get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 20:45:28 +02:00
Niels Lohmann
632a5812a8 Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041) (#5272)
* Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041)

NLOHMANN_DEFINE_TYPE_INTRUSIVE(Type) and its 11 sibling macros produced
broken code for types with no members to serialize. Invoking a variadic
macro so __VA_ARGS__ is empty is only standard-conforming since C++20,
so a plain __VA_OPT__ fix (as tried in #5142) breaks every pre-C++20
build under -pedantic. Instead, make all 12 macros purely variadic and
dispatch on argument count using a sentinel-padded extension of the
existing NLOHMANN_JSON_GET_MACRO idiom, giving full C++11-C++26 support
with no feature-test gate.

Verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros: zero regressions in the existing
unit-udt_macro.cpp suite plus 12 new zero-member test cases.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI failures in zero-member NLOHMANN_DEFINE_TYPE_* macros

Three issues surfaced on PR #5272's real CI that weren't caught by
local testing against a narrower flag set:

- GCC -Werror=noexcept: the four truly-empty from_json bodies (plain
  INTRUSIVE/NON_INTRUSIVE, with and without _WITH_DEFAULT) provably
  never throw but weren't declared noexcept; mark them noexcept
  explicitly. to_json and the derived-type from_json overloads are
  left alone since they genuinely can throw (object assignment /
  delegating to the base class's from_json).
- clang-tidy bugprone-macro-parentheses: false positive on the same
  8 zero-member bodies (Type/BaseType used purely as declarator
  types); suppressed with NOLINTNEXTLINE comments in the same style
  already used elsewhere in this file (see NLOHMANN_JSON_SERIALIZE_ENUM).
- MSVC's traditional preprocessor doesn't fully expand
  NLOHMANN_JSON_CAT(prefix, NLOHMANN_JSON_TYPE_TAG(...))(...) in one
  pass, which broke a pre-existing one-member usage in
  unit-regression2.cpp with syntax errors. Wrap all 12 public
  dispatcher macros in an extra outer NLOHMANN_JSON_EXPAND(...),
  matching the pattern NLOHMANN_JSON_PASTE already uses for the same
  MSVC quirk.

Re-verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros -Wnoexcept, including the exact
files that failed in CI (unit-udt_macro.cpp, unit-regression2.cpp),
against both the modular headers and the re-amalgamated single header.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy misc-const-correctness in unit-udt_macro.cpp

The four zero-member ONLY_SERIALIZE test objects are only ever read
(via to_json), never mutated, so mark them const per clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix derived-type macro dispatch capping members at 62 instead of 63

NLOHMANN_JSON_GET_MACRO resolves 64 positional arguments, with NAME at
position 65. NLOHMANN_JSON_TYPE_TAG dispatches on Type plus the member
list, so it resolves correctly up to the 63 members NLOHMANN_JSON_PASTE
supports. NLOHMANN_JSON_DERIVED_TYPE_TAG dispatched on the two-token
Type,BaseType prefix plus the member list, running out one slot early:
at 63 members, position 65 landed on the last member name instead of a
sentinel and NLOHMANN_JSON_CAT built an undefined identifier such as
NLOHMANN_JSON_DEFINE_DERIVED_TYPE_INTRUSIVE_m63, with the compiler
reporting "unknown type name 'm1'" once per member and nothing pointing
at an argument-count limit.

That silently reduced all six NLOHMANN_DEFINE_DERIVED_TYPE_* macros from
63 members to 62, contradicting the "up to 63 members" contract in
docs/mkdocs/docs/api/macros/nlohmann_define_derived_type.md.

Drop the leading Type and defer to NLOHMANN_JSON_TYPE_TAG so the tag is
computed from BaseType plus the member list, which fits the available
slots. The zero-own-member derived bodies are therefore selected by tag
1 rather than 2, and the sentinel table for the derived tag is no longer
needed.

Add a regression test at the documented maximum for both the plain and
the derived macros; it fails to compile against the previous dispatch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the zero-member macro bodies by intent, not argument count

The dispatch tag was the literal token 1 or N, pasted onto a macro prefix
to select the zero-member or member-carrying body. For the derived-type
macros that reads wrong: their tag is computed after dropping the leading
Type, so the zero-member body was named _1 while taking two parameters
(Type, BaseType).

Emit EMPTY and MEMBERS instead. The mechanism is unchanged -- the tag is
still a token pasted onto the prefix by NLOHMANN_JSON_CAT -- but the body
names now say what they are rather than encoding an argument count that
only lines up for half of the macros.

Collapse the four duplicated zero-member bodies while here: with no
members there is nothing to default, so each _WITH_DEFAULT_EMPTY body was
a byte-for-byte copy of its plain counterpart. They are now one-line
aliases, leaving a single definition of what an empty object serializes
to per intrusive/non-intrusive and base/derived combination.

No functional change: for both zero-member and member-carrying types the
preprocessed to_json/from_json output is token-for-token identical, and
the arity limits are unchanged (63 members, base and derived).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document zero-member support in the macro API reference

docs/mkdocs/docs/features/arbitrary_types.md already gained a note, but
the three api/macros pages are where the parameter contract is actually
specified and they still described member as a non-empty list.

State that the list may be empty on each page, and add a note showing
what the zero-member case generates: an empty JSON object for the plain
macros, and base-type-only serialization for the derived ones. Both notes
record that the WITH_NAMES variants do not support this.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep user macros named EMPTY or MEMBERS out of the member-count dispatch

The dispatch produced the bare token EMPTY or MEMBERS and pasted it onto
the macro prefix afterwards. In between, the token was rescanned, so a
user macro with either name replaced it: with `#define MEMBERS x` in
scope, even NLOHMANN_DEFINE_TYPE_INTRUSIVE(A, member) -- which compiled
before -- expanded to garbage, and `#define EMPTY` broke the zero-member
form.

Paste the suffix onto the prefix directly in the GET_MACRO slot table
instead. Operands of ## are not macro-expanded, so the selected body name
is formed before any user macro can interfere. NLOHMANN_JSON_TYPE_TAG and
NLOHMANN_JSON_DERIVED_TYPE_TAG become NLOHMANN_JSON_TYPE_BODY and
NLOHMANN_JSON_DERIVED_TYPE_BODY, taking the prefix as their first
argument; NLOHMANN_JSON_CAT is no longer needed. The body macro names are
unchanged, and so is the generated code.

Add a regression test that defines EMPTY and MEMBERS around plain and
derived types, with and without members.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Test for EMPTY and MEMBERS so -Wunused-macros accepts them

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 20:44:30 +02:00
Niels Lohmann
ec4bdc398a Document the benchmarks, make them build and compare versions again, and pin Google Benchmark (#5556)
* Document the benchmarks, and make them build and compare versions again

The benchmark project hasn't configured since #4793: download_test_data.cmake
compiles cmake/detect_libcpp_version.cpp relative to CMAKE_SOURCE_DIR, which
is tests/benchmarks when that is the top-level project, so try_run fails and
so does `make run_benchmarks`. The path is now relative to the module itself,
which is the same file for the main build.

The Dump benchmark discarded dump()'s result, which is [[nodiscard]] by now;
it warned, and left the optimizer free to shorten the loop. The result is
now kept with benchmark::DoNotOptimize.

A new cache variable, JSON_BENCHMARK_INCLUDE_DIR, names the directory holding
the nlohmann/json.hpp to benchmark (single_include by default, as before),
so the same benchmarks can be built against two versions and compared.

tests/benchmarks/README.md documents what is measured, how to build and run
the benchmarks, how to read the output, and how to compare two versions with
Google Benchmark's compare.py; it recommends doing so by hand before a
release rather than in CI.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point ci_benchmarks at tests/benchmarks

The target has configured ${PROJECT_SOURCE_DIR}/benchmarks since it was
added in #2561, but the benchmarks live in tests/benchmarks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Pin Google Benchmark to release 1.9.5

The benchmarks fetched Google Benchmark's main branch, so two builds on
different days could measure with different library code, and CMake 3.30
and later warn that the single-argument FetchContent_Populate() is
deprecated. Fetch the 1.9.5 release archive, verified by its SHA-256,
with FetchContent_MakeAvailable() instead. That needs CMake 3.14; Google
Benchmark itself already needed 3.13.

Its -Werror is switched off, so a newer compiler's new warnings cannot
break the pinned release, and its install rules are no longer added.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 20:43:34 +02:00