Commit Graph

1151 Commits

Author SHA1 Message Date
Niels Lohmann
d3ba92d4fb Merge branch 'json-view/19-edit-set' into json-view/20-edit-structure
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:07 +02:00
Niels Lohmann
89d6a7c2b5 Merge branch 'json-view/18-view-object-index' into json-view/19-edit-set
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:28:57 +02:00
Niels Lohmann
95c8d2aa46 Merge branch 'json-view/16-view-simd' into json-view/18-view-object-index
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:28:52 +02:00
Niels Lohmann
d38f5f111f Merge branch 'json-view/15-view-bench' into json-view/16-view-simd
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:28:48 +02:00
Niels Lohmann
d9e4155ba1 Merge branch 'json-view/14-view-compare' into json-view/14b-view-float-layout
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:39 +02:00
Niels Lohmann
7130880754 Merge branch 'json-view/13-view-dump' into json-view/14-view-compare
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:30 +02:00
Niels Lohmann
89bf08760f Merge branch 'json-view/12-view-values' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:23 +02:00
Niels Lohmann
b1595c1b40 Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:13 +02:00
Niels Lohmann
3d6d610fdc Merge branch 'json-view/10-view-document' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:03 +02:00
Niels Lohmann
134b2f0efe Mark json_view.hpp's read() and strlen as Flawfinder false positives
json_document::read is a member function, not POSIX read(), and the C
string overload requires null-terminated input like json::parse.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:01 +02:00
Niels Lohmann
dddb2d6e43 Merge branch 'json-view/08-view-builder' into json-view/10-view-document
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:26:20 +02:00
Niels Lohmann
b59fc6c902 Mark string_ref's strlen as a Flawfinder false positive
string_ref(const char*) requires a null-terminated string, like
std::string_view's constructor.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:26:18 +02:00
Niels Lohmann
ad248a3290 Merge branch 'json-view/19-edit-set' into json-view/20-edit-structure
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:10 +02:00
Niels Lohmann
db73471df4 Merge branch 'json-view/18-view-object-index' into json-view/19-edit-set
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:08 +02:00
Niels Lohmann
c6ac5c85c2 Merge branch 'json-view/16-view-simd' into json-view/18-view-object-index
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:06 +02:00
Niels Lohmann
0e57b3fb28 Merge branch 'json-view/15-view-bench' into json-view/16-view-simd
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:04 +02:00
Niels Lohmann
94c518f94d Merge branch 'json-view/14-view-compare' into json-view/14b-view-float-layout
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:59 +02:00
Niels Lohmann
76ce7e2c84 Merge branch 'json-view/13-view-dump' into json-view/14-view-compare
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:57 +02:00
Niels Lohmann
5277335a9e Merge branch 'json-view/12-view-values' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:55 +02:00
Niels Lohmann
05b6cd0892 Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:53 +02:00
Niels Lohmann
95d10dab70 Merge branch 'json-view/10-view-document' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:51 +02:00
Niels Lohmann
371a8a3d9f Merge branch 'json-view/08-view-builder' into json-view/10-view-document
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	Makefile
#	cmake/ci.cmake
2026-10-01 10:19:49 +02:00
Niels Lohmann
cb51f80e34 Merge branch 'json-view/04-unicode-escapes' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann
ae01d57694 Merge branch 'json-view/03-string-scan' into json-view/04-unicode-escapes
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann
6a757ca675 Merge branch 'json-view/02b-float-parser' into json-view/03-string-scan
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann
9d44e3f359 Merge branch 'develop' into json-view/02b-float-parser
Conflicted only in tests/src/unit-class_lexer.cpp, where develop's #5737
lint fix (CAPTURE(x); -> CAPTURE(x)) collided with this PR's rewrite of
the Eisel-Lemire float tests; kept the PR's new tests and applied the
lint-fixed CAPTURE style. single_include regenerated via make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:53:02 +02:00
Niels Lohmann
e5a89d671f Fix lint debt: enum-macro NOLINTs, doctest as SYSTEM, no-op analyzer (#5737)
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw

The comment said the size_type overflow check in array_index() is only
triggered on special platforms like 32-bit, and the throw was excluded
from coverage. On 64-bit platforms the check is true for SIZE_MAX
itself, and unit-json_pointer.cpp has asserted that case four times
since #5395, so the line is executed in the coverage job. Reword the
comment and remove the exclusion marker so the coverage report notices
if the tests stop reaching it.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name all three C-array check aliases in the enum-macro NOLINTs

NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning
under modernize-avoid-c-arrays only, but clang-tidy emits the same
diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and
hicpp-avoid-c-arrays too. Any user running those checks got a false
positive at every macro expansion, and our own tests needed a local
NOLINT at each call site to work around it.

Name all three aliases in the four macro comments instead, and drop
the now-redundant c-array names from the five test call-site NOLINTs.
Comment-only change; behavior, the public API, and the ABI do not
change. Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Include doctest as a SYSTEM directory instead of disabling warnings for all tests

test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile
options for every non-MSVC compiler, so they were applied to every
translation unit, library headers included, and silenced the CI
warnings meant to check the library's own -Wfloat-equal pragmas. The
only code that actually needed the suppression was the vendored
doctest.h, which was included as a normal (non-SYSTEM) directory.

Include thirdparty/doctest as SYSTEM for test_main, matching what
tests/abi/CMakeLists.txt already does, and drop the two suppressions
from both targets. Verified locally that unit-comparison,
unit-conversions and unit-constructor1 compile clean with
-Werror -Weverything and doctest as -isystem, and that CMake still
configures with JSON_BuildTests=ON.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the no-op ci_clang_analyze target

ci_clang_analyze configured the build with the real compiler and only
then wrapped ninja with scan-build. scan-build intercepts compiles by
overriding CC/CXX, but build.ninja already had the compiler path baked
in from the configure step, so every run bypassed the analyzer: CI
logs show "No bugs found" after a normal build, never an analysis.
The job also used Debian's frozen clang-tools-14 rather than the
image's own clang, and CLANG_ANALYZER_CHECKS still named three
valist.* checkers that current clang merged into security.VAList.

ci_clang_tidy already runs every clang-analyzer-* check (via
.clang-tidy's "Checks: '*'") with warnings as errors, so nothing is
lost by removing the dead job. Delete ci_clang_analyze,
CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from
cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang
matrix, and drop the now-unused clang-tools apt package (iwyu stays
for ci_single_binaries). Reword quality_assurance.md and
assurance_case.md, which described the dead job as a working control,
to say the Clang Static Analyzer checks run through clang-tidy.

Verified that `cmake -DJSON_CI=ON` still configures cleanly and that
ci_clang_analyze no longer appears in the generated build or in any
CMake/workflow file.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable portability-template-virtual-member-function; remove redundant forwards

.clang-tidy disabled three checks "to get the CI going" (#4489,
2024-11-13): portability-template-virtual-member-function,
bugprone-use-after-move and its alias hicpp-invalid-access-moved.

portability-template-virtual-member-function only flagged
output_stream_adapter::write_character/write_characters; annotate
both with NOLINT and re-enable the check.

bugprone-use-after-move flagged several double forwards that have no
effect at runtime:
- from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside
  pack expansions; at() has no ref-qualified overloads and always
  returns an lvalue reference, so the forward is a no-op. Replace with
  plain j.at(Idx) in all four places.
- the move constructor forwards the whole object to its base class
  and then reads other's members. That is item 9 of #5724 (together
  with its cppcheck suppressions) and is left to that change.
- input_adapters.hpp forwards the container twice on purpose, so the
  begin/end iterator types match adapter_type; annotate with NOLINT
  and a comment instead of changing behavior.

The check still flags the move constructor (see above) and two sites
in at(KeyType&&) (both overloads, json.hpp, in the throw's
string_t(std::forward<KeyType>(key)) after
find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk,
so bugprone-use-after-move (and hicpp-invalid-access-moved)
stay disabled for now, with a comment explaining why; re-enable them
once #5689 and the #5724 move-constructor change have landed.

Also resolve the portability-avoid-pragma-once TODO: single_include
never has #pragma once (amalgamate.py strips it) and every supported
compiler accepts it in include/, so keep it disabled with an
explanatory comment instead of a TODO. Fix the stale "json.hpp,
around line 1265" comment in unit-class_parser.cpp, which now points
at the move constructor's actual line.

Behavior, the public API and the ABI do not change. Verified with
clang-tidy 22.1.8 that portability-template-virtual-member-function
now reports nothing, that bugprone-use-after-move/
hicpp-invalid-access-moved report only the known at(KeyType&&) and
move-constructor sites, and that unit-custom-base-class, unit-constructor1,
unit-conversions, unit-element_access2, unit-class_parser and
unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile
under ASan/UBSan and pass with the same assertion counts as before.
Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale and malformed NOLINT comments

json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two
JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a
diagnostic tag under WarningsAsErrors, not a check name, and every
other JSON_ASSERT(false) omits it.

unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)"
comments on "json j = ...;" declarations that are all used with
non-const members afterwards, so the check has nothing to report
there.

unit-constructor2.cpp used a blanket "// NOLINT: access after move is
OK here" on a use-after-move that hides every check on the line;
naming bugprone-use-after-move and hicpp-invalid-access-moved keeps
the intent once those checks are re-enabled (#5724).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 10

* Remove stale .clang-tidy entries

-google-runtime-references disabled a check that neither clang-tidy
22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was
removed upstream. The commented-out HeaderFilterRegex line has been
unused since the active HeaderFilterRegex was introduced in #2561
(2021).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 11

* Remove the GCC C++20 -Wignored-attributes pragma in json.hpp

The pragma (added in #5164) claimed to work around the C++ modules
redefinition errors of #5103, but #5103 is about hard errors (e.g.
"redefinition of std::__is_constant_evaluated()", conflicting
std::integral_constant) that ignoring a warning cannot suppress; they
are traced to GCC PR 124430 and reproduce with <map> or <string>
instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build
following #5103's repro steps still fails with the pragma in place,
and a build of all test TUs with GCC_CXXFLAGS (which enable
-Wignored-attributes) and the pragma removed produces no such
warning. The block only hid a warning class from GCC C++20 users
while suggesting #5103 was handled.

Overlaps #5610, whose hunks touch the closing half of this pragma to
insert the json_literals.hpp include.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 9

* Fix stale doxygen comments hidden by the -Wdocumentation pragma

macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command
for the whole library, which also hides genuine documentation mistakes:

- detail::unescape() documented "@return unescaped string" but returns
  void and unescapes its argument in place; reworded to
  "@param[in,out] s string to unescape in place" and dropped the
  bogus @return.
- basic_json::get()'s copy-conversion overload wrote "converted to
  @tparam ValueType" inside @return, which Doxygen and Clang parse as
  a second, malformed @tparam; changed to "@a ValueType", matching the
  two other get() overloads a few lines above that already use it.

This narrows the gap the -Wdocumentation pragma needs to cover; fully
replacing the Doxygen-only commands it also hides (item 2c) is left
for after #5267.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 2

* Fix -Wextra-semi-stmt at its actual source, not assert()

clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(),
but assert() expands to an expression under glibc and libc++ and does
not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT
with "{if (!(x)) ++assert_counter; }", a bare block followed by a
semicolon at every JSON_ASSERT(...) call site in the library; that
was the actual source of 151 of the 208 -Wextra-semi-stmt sites found
in a Clang 22 -Weverything sweep of the test suite with the flag
removed. Switched to the standard do/while(false) macro idiom, which
does not expand to a statement-plus-semicolon, and corrected the
comment to name the remaining source instead: vendored Doctest's
CAPTURE(x) shim, which already ends in a semicolon.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that unit-assert_macro.cpp now compiles without any
-Wextra-semi-stmt diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites)

* Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt

doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with
a trailing semicolon baked into the macro), specifically so call sites
do not need to add one themselves; most of the ~267 call sites already
follow that convention. The remaining 64 call sites across 20 files
wrote "CAPTURE(x);" anyway, turning into a statement plus an empty
statement and triggering -Wextra-semi-stmt. Dropped the redundant
semicolon at each of those sites.

With item 6 having already made vendored Doctest a SYSTEM include, and
this the last known source of -Wextra-semi-stmt findings, removed the
flag from clang_flags.cmake entirely.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that all 20 touched files, plus a file with no CAPTURE() use
(unit-json_pointer.cpp), compile without any -Wextra-semi-stmt
diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 2 of 2)

* Switch ci_static_analysis_clang off the frozen LLVM 22 dev image

ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries
job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports
"clang version 22.0.0 (...+20251015...)", a pre-release snapshot from
before the LLVM 22 release; the maintainer now updates dev-unstable,
22, and latest instead. Switched to silkeh/clang:22, matching the
other clang jobs on :latest.

Verified with clang-tidy 22.1.8 (the image's actual version) against
this repository's .clang-tidy and library headers what the release
image newly reports compared to :dev:

- readability-redundant-typename fires at ~250 sites across the
  _cpp20-relevant conversion/to_chars headers; the library targets
  C++11 and keeps the typenames, so the check is disabled in
  .clang-tidy, matching how the file already handles checks that
  don't fit a C++11 codebase.
- misc-anonymous-namespace-in-header fires on the two anonymous
  namespaces in from_json.hpp and to_json.hpp; added the alias to
  their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces,
  google-build-namespaces).
- bugprone-std-namespace-modification fires on every addition to
  namespace std: the std::hash, std::formatter and std::swap
  overloads in json.hpp, and the std::tuple_size/std::tuple_element
  specializations in iteration_proxy.hpp (this last file is not named
  in #5725's item 5, found by actually running clang-tidy 22.1.8
  against the current tree). All six are legal, deliberate additions
  to namespace std (explicit/partial specializations of std types, or
  the pre-C++20 std::swap overload); annotated each with the check
  name next to its existing cert-dcl58-cpp NOLINT.
- modernize-avoid-c-style-cast reported nothing new.

Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the
find_program search lists in ci.cmake so a local "maximal warnings"
configure prefers the current toolchain version over an older one on
PATH.

#5725 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Regenerate cmake/gcc_flags.cmake for GCC 16.2.0

GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and
ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0,
so the "maximal warnings" job was missing warnings introduced since
15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops.

Regenerated with https://github.com/nlohmann/gcc_flags (patched
locally to not crash on an option whose "-x c++ <opt> -" probe fails
before it reads stdin, e.g. -Wabi=; the tool otherwise raises
BrokenPipeError instead of recording the option as an error) run
against g++ 16.2.0 in the official gcc:16 Docker image, keeping the
documented -Wno-* exclusions and the same alphabetical placement
scheme as before.

Also added three GCC 16 warnings the generator cannot discover on its
own because it only probes value ranges/lists it finds in the -Q
option name itself, not in the enum choices --help=warnings documents
separately:

- -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified
  these compile cleanly with g++ 16.2.0.
- -Wstrict-flex-arrays: deliberately NOT added, unlike the other two.
  Without -fstrict-flex-arrays (which the library does not enable, as
  it would change codegen for flexible array members), GCC prints
  "'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is
  not present" on every translation unit, and under our -Werror that
  note itself aborts the build. This differs from the harmless
  no-op warnings already kept in the file (-Whsa, -Wsynth,
  -Wunreachable-code, -Wunsafe-loop-optimizations), which emit
  nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the
  flags GCC 16 adds, but did not anticipate this failure mode.

Verified: compiled the library header and a representative set of
test translation units (including ones touched by items 1, 3, 8, 9,
10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror
under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings;
ran the full local test suite (129/129 passing, unrelated to this
compiler) as a regression check. CI must still confirm the actual
ci_test_gcc / ci_test_standards_gcc targets end to end, since this was
verified with direct g++ invocations rather than through the CMake/
CXXFLAGS environment-variable plumbing in ci.cmake.

#5725 item 7

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid std::basic_string<CharType> for non-character output_adapter CharType

output_adapter<CharType, StringType> defaulted StringType to
std::basic_string<CharType>, and (with JSON_NO_IO undefined) always
declared a std::basic_ostream<CharType>&-taking constructor. For
CharType with no non-deprecated std::char_traits specialization (only
std::uint8_t is ever used this way, by the binary writers), simply
naming either type - as an unused default template argument, or as an
unused, never-called constructor's parameter type - instantiates
std::char_traits<CharType> merely to name it, which some standard
libraries mark deprecated: with the library-wide -Wdocumentation
pragma (item 2's other half, left for a later commit) temporarily
removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec)
with std::vector<std::uint8_t>& got one -Wdeprecated-declarations
warning per binary writer at the old output_adapters.hpp:193.

Replaced the eager std::basic_string<CharType> / std::basic_ostream
<CharType> defaults with a bool-tagged partial specialization (not
std::conditional, which requires naming both branches' types up
front regardless of which is selected, reproducing the same warning)
that only ever names std::basic_string<CharType> / std::basic_ostream
<CharType> when CharType is actually one of char, wchar_t, char16_t,
char32_t, or (with __cpp_lib_char8_t) char8_t. For any other
CharType, output_adapter's StringType and ostream-constructor
parameter fall back to two distinct empty placeholder types, kept
distinct so the two constructor overloads do not collide into a
single redeclaration.

Public API / behavior: passing a std::basic_string<std::uint8_t>& or
std::basic_ostream<std::uint8_t>& directly to a binary writer's
output_adapter now fails to compile instead of compiling with a
deprecation warning; this was neither documented nor tested. All
documented uses (std::vector<CharType>, std::basic_ostream<CharType>
and StringType for character CharType) are unaffected.

Verified with Apple clang 21 / libc++, with the two -Wdocumentation*
"ignored" pragma lines in macro_scope.hpp temporarily removed and
-std=c++11/c++20 plus the project's -Weverything flag set: calling
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a
std::vector<std::uint8_t> now produces no char_traits<unsigned char>
(or any other) deprecation warning, while the char-based string- and
ostream-adapter paths, and a to_cbor/from_cbor round trip, still
compile and run correctly; also verified with GCC 16.2.0. Ran the
full local test suite, including the binary-format unit tests
(unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson,
unit-bon8, unit-binary_writer_sinks, unit-binary_formats,
unit-custom-binary-type): 129/129 passing.

#5725 item 2 (step a)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the library-wide -Wdocumentation pragma; fix what it hid

macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang
diagnostic region over the entire library that ignored -Wdocumentation
and -Wdocumentation-unknown-command. Removed both pragmas and fixed
every finding a full -Wdocumentation (which implies
-Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync)
build reports, so the library now compiles clean under Clang's
documentation checks without a blanket suppression. Overlaps #5267,
which is still open and edits a nearby doc block (json.hpp's
get()/get_impl() @return, already fixed in the item 2 step (b) commit
of this branch); this commit does not touch that block again.

Unknown Doxygen alias commands (Doxyfile removed in #3071, so these
were never rendered by anything) rewritten as plain prose, keeping the
same information:
- @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp,
  json_reverse_iterator.hpp): now "This class satisfies the following
  concept requirements (REQ-JSON-0N):".
- @liveexample{prose,example-id} (three sites in json.hpp): kept the
  prose, dropped the command wrapper and the trailing example-id
  (docs/mkdocs/docs/examples/*.cpp still exist and are used directly
  by the rendered docs, not through this in-header alias) and
  unescaped the "\," commas that were only needed for the old alias's
  comma-separated argument syntax.
- @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1):
  now "Complexity: X".

Backslash sequences Clang's comment lexer tried to parse as commands,
escaped to render as literal backslashes:
- lexer.hpp get_codepoint(): two `\u` occurrences.
- binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two
  `\x00` occurrences.
- serializer.hpp: three `\uXXXX` occurrences (constructor @param,
  append_codepoint_to_string_buffer() @brief, and the ensure_ascii
  member comment).

One finding remained after all of the above: Clang reports
"declaration is marked with '@deprecated' command but does not have a
deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...)
overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to
__attribute__((deprecated(...))) for Clang. Several isolated
reproductions of this exact declaration shape - doc comment,
template<>, two stacked __attribute__ macros, an overload set sharing
the name - did not reproduce the warning, so this looks like a
Clang comment/declaration-association quirk specific to this overload
inside the much larger basic_json class template, not an actual
documentation defect. Rather than keep the pragma library-wide for one
Clang false positive, added a tightly scoped
-Wdocumentation-deprecated-sync push/pop around just that overload.

Verified with Apple clang 21 and the project's actual -Weverything
flag set (cmake/clang_flags.cmake) on the full header at -std=c++11
and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled
clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so
this only confirms no unrelated breakage). Ran make check-amalgamation
and the full local test suite: 129/129 passing.

#5725 item 2 (step c)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Take the JSON value by const reference in the array and tuple from_json paths

Review feedback on #5737 (gregmarr): once the no-op std::forward calls
are gone, the forwarding references have no purpose. from_json_fn
passes the value as const BasicJsonType&, so these functions were only
ever instantiated with a const lvalue anyway.

The std::array, std::pair and std::tuple overloads of from_json and
their helpers now take const BasicJsonType& and pass j on unchanged.
Because the deduced BasicJsonType is now the plain type, tuple_type and
the static_assert name const BasicJsonType& explicitly, so the
reference checks are unchanged: get<std::tuple<const std::string&>>()
still works, and get<std::tuple<std::string&>>() still fails the same
static_assert. from_json_tuple_get_impl keeps its forwarding reference,
since tuple_type calls it through std::declval.

Behavior, the public API and the ABI do not change. unit-conversions,
unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3,
unit-deserialization, unit-noexcept, unit-items, unit-allocator,
unit-custom-object-type, unit-ordered_json2 and
unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with
unchanged assertion counts. Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:37:47 +02:00
Niels Lohmann
cff0a61369 Share DOM SAX position handling; fix stale parser and lexer comments (#5731)
* Share the diagnostic-position setter of the DOM SAX parsers

json_sax_dom_parser and json_sax_dom_callback_parser each had a private
copy of handle_diagnostic_positions_for_json_value(), identical except
for comments. Move the body into one static member function,
detail::diagnostic_positions::set_from_lexer(value, lexer), which both
classes call with their lexer pointer. basic_json befriends the new
struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position
members are private.

The discarded case is reached through the callback parser, so the
LCOV_EXCL markers that only the dom parser's copy had are gone. The
NOLINT on the unreachable default case loses the stray
"-warnings-as-errors", which is not a check name.

The start-position setup in start_object()/start_array() is left alone,
as #5706 is editing the callback parser's versions.

Behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the parser comments on recursion and skip_to_state_evaluation

The class documentation called the parser a recursive descent parser,
but sax_parse_internal() is a loop that keeps the open containers on an
explicit stack. The comment at the end of an array and of an object
said the flag is set to false while the code below it sets it to true.
Describe what the code does instead.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update the discard_number_values comments to the current number path

The comments explaining the accept() shortcut in convert_number() and
the member documentation still argued in terms of strtoull()/strtoll()
and errno, which #5283 replaced with convert_integer(), and pointed at
scan_number() instead of convert_number(). They also did not say that
scan_number_bulk_contiguous() converts integers itself, so the shortcut
is only reached for input without bulk access, with
JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back.

Rewrite both comments to describe the digit-count check in front of
convert_integer(), keeping the 18-digit bound and the json_sax_acceptor
argument. The stale <cstdlib> comment is left for after #5616, which
edits that include block.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* List the UTF-8 validators instead of calling the DFA the only one

The documentation of decode() called the Hoehrmann DFA the single
source of truth for UTF-8 validation. It is used only by the serializer
and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text
strings). The lexer's scan_string() switch, validate_one_utf8() /
valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer)
and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges
on their own.

Replace the sentence with a list of the four validators, what each is
used for, and a note that they must accept the same sequences. Sharing
code between them was considered and dropped: it would save a few lines
in a validator that is entangled with BON8 pushback, and #5677 is
editing the BON8 byte path.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale doc comments and include lists in the input headers

input_adapters.hpp included <memory> and <numeric> for the removed
shared_ptr-based adapter design but used neither; it called
(std::min) without including <algorithm>. json_sax.hpp used
std::numeric_limits without including <limits>. Also corrected
comments that no longer matched the code: input_stream_adapter does
not skip the input's BOM (the lexer's skip_bom() does), the
span_input_adapter comment named the no-longer-existing
input_buffer_adapter type, lexer::get_string() does not reset the
token, binary_reader's get_number() doc opened with /* instead of
/*! (so Doxygen skipped it) and omitted BON8 from its endianness
note, and the UBJSON-binary-types note did not mention that BJData
'B' arrays are read as binary.

Left out: the lgtm suppression on lexer.hpp's scan_number() (in
#5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's
start_object/start_array docs (in draft #5267's hunk), per the
verdict's conflict list.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the strict-EOF/release_lookahead/error block in parser::parse()

json_sax_dom_callback_parser and json_sax_dom_parser branches of
parser::parse() ran the same ~25 lines after sax_parse_internal():
the strict-mode EOF check (raising parse_error.101 through the SAX
parser), release_lookahead() in non-strict mode, and mapping an
errored SAX parser to a discarded result. The two copies had already
drifted apart in formatting and in the second copy's "see above"
comment.

Add a private parse_dom(DomSax&, strict) member that runs this shared
sequence once and returns whether the SAX parser did not error; both
branches of parse() now only construct their DOM SAX parser, call
parse_dom(), and (for the callback parser) map a discarded top-level
value to null. sax_parse() is left untouched, since it only runs the
EOF check and release_lookahead() when sax_parse_internal() succeeded,
unlike parse(), which runs them unconditionally.

Behavior-preserving: same operations in the same order for both SAX
parser kinds. Verified with unit-class_parser (strict/non-strict,
callback and non-callback), unit-deserialization and
unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean
make amalgamate / make check-amalgamation diff.

Overlaps #5601, which touches the same lines.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 2

* Share the code point to UTF-8 encoding between the wide-string helpers and the lexer

The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three
times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32
code point, in the UTF-16 helper for both a BMP code unit and a valid
surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The
copies had drifted: the UTF-32 helper masked the leading bits of each
byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift
alone, even though both give the same result for a code point that is
already known to be in range.

Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder
that invokes a callable once per output byte, most significant byte
first. Use it in the three valid-code-point branches (UTF-32 code
points up to U+10FFFF, UTF-16 code units outside the surrogate range,
and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where
out forwards to add(). The UTF-16 helper's deliberate pass-through of
malformed surrogate units and the UTF-32 helper's 0xFF sentinel for
code points above U+10FFFF are untouched, since neither reaches the new
helper.

Behavior-preserving: same bytes in the same order for every valid code
point, verified with unit-class_lexer, unit-class_parser,
unit-deserialization, unit-wstring and the non-test-data parts of
unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse
microbenchmark that shows no change (about 73 ms either way, median of
3, 1M escape sequences). single_include/ regenerated with make
amalgamate; make check-amalgamation leaves a clean tree.

Overlaps #5704, which rewrites the wide_string_input_helper
specializations touched here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 6

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:32:32 +02:00
Niels Lohmann
3926fcaac3 Deduplicate serializer dump code; fix stale includes, docs, and lint (#5729)
* Share scalar serialization between dump_internal and dump_value

dump_value()'s cases for string, binary, boolean, number_integer,
number_unsigned, number_float, discarded and null were a byte-for-byte
copy of dump_internal()'s (added together in #5285 for the iterative
fallback path). Any future change to scalar output had to be made in
both places, or the recursive and depth-limited paths would silently
start producing different bytes.

Extract the shared cases into a private dump_scalar() and have both
dump_internal() and dump_value() call it. Output is unchanged: dump(),
dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t
variants are byte-identical over the json_test_data corpus before and
after, and dump() throughput on a scalar-heavy document is unaffected.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop serializer.hpp's dependency on binary_writer.hpp

The only use of binary_writer in serializer.hpp was
binary_writer<BasicJsonType, char>::to_char_type() to write the
U+FFFD replacement character's three bytes. With CharType=char this
is an identity conversion, so the include of binary_writer.hpp (and
transitively binary_reader.hpp) pulled in a large, unrelated header
for a no-op call.

Write the three bytes directly instead. serializer.hpp compiles
standalone with -Wall -Wextra -Werror, with and without
-funsigned-char, and unit-serialization's error_handler_t::replace
cases (with and without ensure_ascii) still pass. Moving
binary_writer's to_char_type/to_msgpack_length to its private section
is left as an optional follow-up.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale #include lines in the output headers

output_adapters.hpp included <algorithm> and <iterator> for std::copy
and std::back_inserter, which have not been used there since #3569
(2022). serializer.hpp included <algorithm> for std::reverse (also
unused), <cmath> for labs/isnan/signbit (only std::isfinite is used)
and <utility> for std::move (nothing from <utility> is used there),
while using std::next without including <iterator> at all, relying on
getting it transitively through output_adapters.hpp's own stale
<iterator>.

Drop the unused includes, add <iterator> for std::next, and correct
the remaining include comments. Both headers still compile standalone
with -Wall -Wextra -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides

output_vector_adapter, output_stream_adapter and output_string_adapter
declared their write_characters(const CharType*, std::size_t) override
JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with
a null pointer and length 0 for an empty string or binary value; the
type-erased call path only stayed silent under UBSan because the
static callee at those call sites is the unattributed virtual base.
A nonnull attribute on a definition lets GCC and Clang assume the
parameter is non-null inside the function body even when the call is
virtual, so this was latent undefined behavior, not just style.

Drop the attribute from the three overrides and document the
(nullptr, 0) contract on output_adapter_protocol::write_characters.
unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise
empty binary/string payloads through the stream and vector/string
adapters) pass under -fsanitize=address,undefined,nonnull-attribute.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update stale serializer doc comments to match the current implementation

dump_internal()'s doc block still described the pre-#5285/#5449
implementation: an escape_string() function that does not exist
(the function is dump_escaped), integer conversion "implicitly via
operator<<" (dump_integer actually uses a digit-pair lookup table),
and floating-point conversion via "%g" (IEEE-754 types go through
to_chars, others through snprintf). dump_value()'s comment said
elements are pushed for dump_internal to walk, but it is
dump_iteratively() that walks the stack. dump_escaped(), dump_integer()
and dump_float() each said they write "to output stream @a o", which
has not been true since the writer moved to write_buffer.

Doc-only change; no behavior, API or ABI impact.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge duplicate byte-to-hex helper and drop stale '| 0' promotions

serializer::hex_bytes() and binary_writer::hex_byte() had identical
bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it
from both. Also drop the `| 0` at the two serializer call sites
(hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8)
added it so that `ss << std::hex << (byte | 0)` printed a number
rather than a char with the old stringstream writer; the int result
just narrows back to uint8_t now, so it was a no-op.

Behavior is unchanged: unit-serialization, unit-bon8 (whose
type_error.316 messages exercise this code) and unit-diagnostics
pass, and both headers still compile standalone with -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Trim two stale lint suppressions in serializer.hpp

dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried
NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg,
left over from the snprintf-based implementation (#3088); there is no
variadic call on that line, so keep only the qualified-auto
suppressions it actually needs. remove_sign()'s assert checked
`x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a
NOLINT(misc-redundant-expression) to hide it; the second conjunct is
always true once x < 0, and has been since 6ce2f35ba (2019), so
reduce the assert to `x < 0` and drop the suppression instead of
masking it.

Both are documentation-only changes to assertions/suppressions, not
behavior. The to_chars.hpp `#if 0` branch this item also flagged is
left alone, next to draft PR #5634's pending hunk.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the Hoehrmann SPDX copyright line to string_utils.hpp

serializer.hpp carried the SPDX-FileCopyrightText line for Björn
Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d
table) has lived in string_utils.hpp since #5185 (d19f7f5dc);
serializer.hpp now only calls decode(). Move the copyright line to
where the code it covers actually is.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop calling std::localeconv() on every dump()

The serializer constructor snapshotted std::localeconv() into a
locale_chars member on every dump(), even though the only reader is
dump_float(number_float_t, std::false_type)'s snprintf path, taken
only for a number_float_t that is neither IEEE single nor double.
localeconv() is not required to be thread-safe with setlocale(), so
every dump() paid for and raced on a lookup that almost never mattered.

Remove locale_chars and the locale member. Right before the
thousands-separator/decimal-point fixups in the snprintf path, read
std::localeconv() into local thousands_sep/decimal_point variables
(null-checked, first byte only, as before) - the same way
lexer::get_decimal_point() already does since #5597. Output is
unchanged unless the locale changes during a single dump(); in that
case the fixups now match what snprintf just produced, instead of a
value snapshotted before the call.

Overlaps draft PR #5608, which touches the same constructor and
dump_float() lines to move this code into a new
dump_float_snprintf(); this lands the lookup change now as #5709 asks,
and #5608 can do the lookup inside dump_float_snprintf() when it
rebases.

Verification: the full json_test_data corpus (742 files, dump(),
dump(4) and dump(-1,' ',true)) is byte-identical to before the change
under the C locale. Added a test pinning the new per-conversion
lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that
switches on its first write, after the serializer's write buffer has
been flushed once but before a later float is converted) and checks
the decimal point is still normalized using the locale active at
conversion time. On a platform where long double is IEEE-754 double
(e.g. 64-bit Arm), dump_float() takes the locale-independent
to_chars() path and the test is a no-op there; it is meaningful on a
platform where long double is extended precision (most x86 targets).

Part of #5709 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:31:42 +02:00
Niels Lohmann
18dd5663b0 Diff deeply nested values without recursing per nesting level (#5548)
* Diff deeply nested values without recursing per nesting level

diff() descended into both values once per nesting level, and compared
them with operator== on every level on the way, which recurses as well.
Values nested deeply enough - 25,000 levels on an 8 MiB stack - exhausted
the call stack and terminated the process, although parse() accepts
them without complaint. On such a chain the per-level comparisons and
path strings also made diff() quadratic in time and memory.

Both the recursion and operator== only descend as far as the source is
nested. So diff() first checks, recursing at most diff_depth_limit()
(128) levels, whether the source is nested more deeply than that. If not
- all but a vanishing minority of values - the recursive algorithm
diffs it exactly as before, now as diff_recursively(). Otherwise
diff_iteratively() walks the two values on an explicit stack, emitting
the same operations in the same order. It does not compare arrays and
objects with operator== up front (equal ones yield no operations
anyway), keeps the path in one buffer instead of a new string per
level, and hands every subtree that is not nested too deeply back to
diff_recursively(), so equal parts are still skipped quickly.

The check costs one pass over the source. On a 3,000-object document
that is about 30% of diffing two equal values (which is just an
operator== call), about 10% of diffing values that differ in a few
places, and noise when arrays change length. Once operator== no longer
recurses (#5390), the check can go.

Tests check that the patch reproduces the target at every depth up to
300, for json and ordered_json, including reordered members. They also
check the exact operation for a difference deep inside, and diff values
nested 100,000 levels deep.

Fixes #5393 for diff().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make diff_frame a member struct that declares its special members

GCC's -Weffc++ (an error in CI) asks a class with pointer members, a
user constructor and a non-trivial destructor to declare its copy
constructor and copy assignment; diff_frame's vector and basic_json
members make its destructor non-trivial. Declare all five as defaulted,
which also satisfies clang-tidy's special-member-functions check. Leave
their exception specifications implicit: GCC 4.8 rejects an explicit
one that differs from the implicit one, as it does for flatten_task in
#5517.

The converting constructor cannot throw, and is now declared noexcept
for GCC's -Wnoexcept, which flags the emplace_back() under C++26
otherwise. The struct also moves from diff_iteratively() into the class,
like dump_frame in the serializer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the shared recursion limit in diff()

diff_depth_limit() is gone in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Diff fewer nesting depths so the test does not time out under Valgrind

Checking every depth up to 300 made test-json_patch exceed the 1500 s ctest
timeout in ci_test_valgrind. Check the depths up to 16, those around the
recursion limit of 128, and 300 instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark the diff frame's value-initialized members for clang-tidy

The braces are kept for GCC's -Weffc++, as in json_sax.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Bound diff()'s descent with a depth count instead of scanning the source

Now that operator== no longer recurses (#5390), diff() can keep its per-level
equality shortcut all the way down. It diffs recursively for the first
detail::recursion_depth_limit() levels, as merge_patch() does, and hands
anything deeper to diff_iteratively(). The nesting_exceeds() scan, which
cost about 30% on equal documents, is gone, and diff() is on par with
develop again.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Note that the diff frame reference is invalidated by pop_back() too

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep diff()'s recursive levels small and its result elided

diff_recursively built every patch operation in place from initializer
lists. Unoptimized builds give each of those temporaries its own stack
slot, so every level of the bounded descent cost kilobytes of stack
(about 6 KB with clang -O0), and the 128 recursive levels overflowed the
1 MB stack of MSVC Debug in the "deeply nested values" test. The
operations and the key comparison of two objects are now built by
separate functions, which diff_iteratively shares, and both diff
functions append to one result instead of returning a patch per level
that the caller copies. With clang -O0, diffing values nested 300 levels
deep now peaks at about 190 KB of stack instead of 880 KB.

Since diff() now owns the only returned value, clang's -Wnrvo no longer
reports the returns of diff_recursively, which alternated between the
local patch and diff_iteratively's result.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Copy the diff frame's members instead of holding a reference to it

The loop in diff_iteratively held a reference to the top frame, which
enter() invalidates when it pushes and the end of the loop invalidates
when it pops. Nothing used it afterwards, but a later change could. As in
the other iterative walks, the members the loop reads are now copied out
as constants and the ones it advances are changed through stack.back().
The frame as a whole is not copied: it holds the common keys and the
"add" operations of an object.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:29:32 +02:00
Niels Lohmann
a32f61eb98 Fix update() and merge_patch() when the argument is *this or one of its members (#5678) 2026-09-30 23:04:57 +02:00
Niels Lohmann
1d675cdb46 Fix CI jobs that check less than they claim; move arm64 to GitHub (#5733)
* Fix the ci_cmake_flags wiring so every option is checked

The CMake 3.31.6 flag list referred to itself before it was defined,
so only JSON_BuildTests was checked with that version. The targets for
the CMake running the build ("_2") were created but never added to
ci_cmake_flags, and the three versions shared one build directory.
JSON_StrictNulHandling was not in the list at all.

Use the 3.5.0 list for 3.31.6, add JSON_StrictNulHandling, and create
one ci_cmake_flag_<flag> target per option for the running CMake with
its own build directory. Also use the function parameter in the
COMMENT, refresh the stale version comment, and let ci_clean remove
the downloaded cmake-<version> directories instead of the long-gone
cmake-3.5.0-Darwin64.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix ci_test_clang_libcxx_cxx* jobs silently building without warnings

CMake only seeds CMAKE_CXX_FLAGS from the CXXFLAGS environment variable
when the cache entry is unset, so the explicit -DCMAKE_CXX_FLAGS="-stdlib=libc++"
argument made it ignore CXXFLAGS="${CLANG_CXXFLAGS}" entirely. The six
ci_test_standards_clang (..., libcxx) jobs therefore compiled without
-Weverything/-Werror while their libstdc++ siblings did use them.

Pass -stdlib=libc++ through the same CXXFLAGS value instead of a separate
-D argument, and give the target its own build directory
(build_clang_libcxx_cxx${CXX_STANDARD}) so it no longer shares a CMake
cache with the libstdc++ variant. Suppress the resulting
-Wthread-safety-negative finding from libc++'s std::mutex annotations,
which fires on doctest's reporters in this translation unit only.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the no-op AppVeyor with_win_header job

The with_win_header matrix entry patched Windows.h into
single_include/nlohmann/json.hpp before building, but JSON_MultipleHeaders
has defaulted to ON since #3532 (2022-06), so CMakeLists.txt points the
tests at include/ and the patched single header is never compiled. The
job has been a no-op VS2015 build since then.

Windows.h coverage already exists through tests/src/unit-windows_h.cpp
(#3631), which runs in every MSVC job. Delete the dead matrix entry and
its before_build steps, and cite unit-windows_h.cpp from the QA page.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove unused ci_oclint and ci_pvs_studio targets

No workflow invokes ci_oclint, ci_pvs_studio, or their tool discovery.
ci_oclint also had a side effect on every JSON_CI configure: it copied
the single header into src_single/all.cpp and added an add_executable()
for it without EXCLUDE_FROM_ALL, so a plain build compiled a 1.2 MB
translation unit that only that unused target consumed. ci_pvs_studio
duplicates the Makefile's pvs_studio target, which is kept.

Also drop the duplicate --check-level=exhaustive flag passed twice to
the same ci_cppcheck invocation.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop Dependabot from proposing astyle bumps

astyle is deliberately pinned at 3.4.13 because newer versions reformat
unrelated lines and this version defines the formatting that
check_amalgamation.yml enforces. Without an ignore rule, Dependabot
keeps opening PRs for every new astyle release (most recently #4580,
#4942, #5445, #5448), each of which fails the amalgamation check and
gets closed unmerged.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move Linux arm64 CI from dead Cirrus CI to ubuntu-24.04-arm

Cirrus CI stopped reporting check runs on develop sometime after
d10879bca (2026-05-26); every commit since has only github-actions
check runs, so .cirrus.yml silently lost its only consumer while
README.md, FILES.md and the QA page kept advertising the coverage.

Add a ci_test_arm64 job to ubuntu.yml using the same pinned
actions/checkout and lukka/get-cmake actions as the other jobs, on the
native ubuntu-24.04-arm runner, with a step that confirms uname -m
reports aarch64. Delete .cirrus.yml and its README badge and FILES.md
section, and update the QA page's arm64 row.

Part of #5715

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make scan-build fail on findings and drop irrelevant checkers (#5715 item 4a)

ci_clang_analyze ran scan-build without --status-bugs, so the job
passed whenever the ninja build succeeded, no matter what the
analyzer found ("No bugs found" in a green run gave no signal either
way). It is also missing --use-analyzer=${CLANG_TOOL}, so scan-build
picks whichever clang happens to be first on PATH inside the
silkeh/clang:dev container instead of the one this file already
selected and versioned.

Add --status-bugs and --use-analyzer=${CLANG_TOOL} to the scan-build
invocation. While here, drop the osx.*, webkit.*, fuchsia.*, and
optin.mpi.* checkers from CLANG_ANALYZER_CHECKS: none of them apply
to this portable C++ library, and leaving them enabled only adds
noise once the job can actually fail on a finding.

The job is currently clean (0 bugs), so this alone does not surface
any new finding; it only makes the existing "no bugs found" result
authoritative. This is 4a of 3 independent steps in #5715 item 4;
4b (Infer) and 4c (IWYU) still need their existing findings triaged
before --fail-on-issue/-Xiwyu --error can be added, and are handled
in separate commits.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the amalgamation/format check's file set and add BUILD.bazel (#5715 item 5)

The amalgamation/format check existed three times with three different
file sets: the Makefile's pretty/check-amalgamation, the pull_request-only
check_amalgamation.yml workflow, and the ci_test_amalgamation CMake target
that also runs on direct pushes to develop/master/release/*. The CMake
target's glob was a strict subset of the workflow's (missing the
docs/mkdocs/docs/examples/*.hpp headers, tests/abi/, tests/cmake_*/project/,
tests/cuda_example/, tests/fmt_formatter/, and tests/module_cpp20/), and it
never checked BUILD.bazel at all, so a misformatted file in any of those
paths, or a stale BUILD.bazel, could reach develop through a direct push
even though the PR-only workflow would have caught it.

Make ci_test_amalgamation glob the same roots (docs/mkdocs/docs/examples,
include, tests) and extensions (*.hpp, *.cpp, *.cu) as check_amalgamation.yml,
excluding tests/thirdparty/ and tests/abi/include/nlohmann/ the same way, and
regenerate and diff BUILD.bazel next to json.hpp/json_fwd.hpp. Also add
docs/mkdocs/docs/examples/*.hpp to the Makefile's pretty/pretty_format
targets, which were missing the four custom_*_type.hpp example headers, and
drop the stale "called by Travis" comment on check-amalgamation (Travis is
gone; nothing currently calls that Makefile target from CI).

Leaves the workflow itself untouched: it deliberately runs amalgamate.py
from a fresh develop checkout so a PR cannot change the tool that checks it.

Overlaps #5610 and #5621, which each add a new amalgamated header and touch
the same INDENT_FILES/ci_test_amalgamation/check_amalgamation.yml hunks.

Verified with `make check-amalgamation` on this branch: clean, no diff.
#5715 item 5.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Download prebuilt CMake binaries on Linux x86_64 instead of building from source (#5715 item 6)

ci_get_cmake() downloaded the source tarball of CMake 3.5.0, 3.31.6, and
4.0.0 and compiled each one completely (including CMake's own test
helpers) with -DCMAKE_POLICY_VERSION_MINIMUM=3.5 as a workaround for
building old CMake with a newer one. On CI this made the
ci_cmake_options (ci_cmake_flags) job take about 11 minutes, most of it
spent building CMake itself, even though Kitware has published
ready-to-run Linux x86_64 archives for all three of these releases
since 3.20 (lowercase platform name).

On Linux x86_64, download and unpack the prebuilt
cmake-<version>-linux-x86_64.tar.gz archive instead and point the
existing ${var} output at its bin/cmake, skipping the configure/build
steps and CMAKE_POLICY_VERSION_MINIMUM entirely. Keep the previous
source build as a fallback for any other platform (macOS, Linux
aarch64), since Kitware does not publish binaries for every
CMake/platform combination this project might build on.

Verify the downloaded archive against Kitware's own published checksum
before unpacking it: download cmake-<version>-SHA-256.txt alongside the
archive and run `sha256sum -c` on the matching line. A CI job that wgets
and untars a binary from a release page with no integrity check is a
supply-chain gap; Kitware has published this file for every release
since 3.20, so checking it costs one extra download and one grep.

As a separate, mechanical change: the ci_cmake_options job's container
only needed to stay on ubuntu:focal for the source build's
libssl-dev dependency and its own aging toolchain; now that the
Linux/x86_64 path never compiles CMake, drop libssl-dev from its apt
install line and move the job to ubuntu:24.04 (Ubuntu 20.04 left
standard support in May 2025). ci_clean already removes the
cmake-3.5.0/cmake-3.31.6/cmake-4.0.0 directories from the #5715 item 3
fix, and the prebuilt path reuses those same directory names, so no
further cleanup changes are needed.

Overlaps #5598, which edits the same ci_cmake_options matrix line in
ubuntu.yml; a rebase may be needed once that lands.

Verified locally: `cmake -S . -B build -DJSON_CI=On` configures cleanly
on macOS/arm64 (source-build fallback branch) and on Linux/x86_64 in an
ubuntu:24.04 Docker container (47 `ci_cmake_flag_*` targets generated,
one built and run successfully); `.github/workflows/ubuntu.yml` still
parses as valid YAML; downloaded the real v3.31.6 Linux x86_64 archive
and SHA-256 file from Kitware and confirmed the `grep | sha256sum -c`
pipeline both accepts the genuine file and is anchored to the exact
filename (not a prefix match).

CI must confirm: the prebuilt-binary path actually runs on the
ubuntu-latest/ubuntu:24.04 x86_64 runner, all `ci_cmake_options`
entries still pass with the new container's GCC, and the job's
runtime drops from roughly 11 minutes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make Infer fail on findings, with a type-level baseline for the ~174 pre-existing ones (#5715 item 4b)

ci_infer ran `infer run` without --fail-on-issue, so the job passed
regardless of what Pulse found; the last recorded run (35829411620,
commit 1054b2097) logged "Found 174 issues" and still went green.
report.txt was also never uploaded, so the full finding list was only
ever visible in the truncated 5-issue console excerpt.

Add a repository-root .inferconfig (auto-discovered by Infer; passing
--project-root on the `infer run` invocation makes sure it is found
even though the analysis runs from build/build_infer) that sets
fail-on-issue and disables the six PULSE issue types that made up all
174 findings in that run: PULSE_UNNECESSARY_COPY_ASSIGNMENT (129),
PULSE_UNNECESSARY_COPY (22), PULSE_UNNECESSARY_COPY_INTERMEDIATE (15),
PULSE_RESOURCE_LEAK (5), PULSE_CONST_REFABLE (2), and
PULSE_UNNECESSARY_COPY_OPTIONAL (1).

This is a deliberate, narrower fix than "triage and fix everything in
this PR": the visible sample is entirely doctest-macro copies in test
code (for example tests/src/unit-algorithms.cpp:141 and
tests/src/unit-bjdata.cpp:3706), but 169 of the 174 findings were never
uploaded anywhere and this PR cannot respectably claim to have fixed
issues it never saw, including the resource-leak and const-refable
ones that are the most likely to be genuine bugs. Disabling by issue
type is a coarser baseline than a per-finding one (Infer has no
built-in per-finding baseline short of the two-run `infer reportdiff`
workflow, which this repository does not have the CI infrastructure
for), but it has the same effect today: the job goes from always green
to green-only-when-clean-of-everything-else, so CI now fails the
moment a *new* issue type appears, and report.txt is uploaded as a
workflow artifact on every run (including failures) so the six
disabled types can be triaged and re-enabled incrementally in follow-up
PRs.

#5715 item 4b. 4a (scan-build) and 4c (IWYU) are handled in separate
commits.

Verified: .inferconfig parses as JSON, ubuntu.yml still parses as
YAML. Infer itself is not available in this environment (v1.3.0 tar.xz
requires a Linux x86_64 runner), so CI must confirm that `infer run
--project-root ... -- make` picks up .inferconfig, that fail-on-issue
takes effect, and that the six disabled types actually suppress the
existing findings without also hiding an unrelated new one.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix IWYU findings for json.hpp/json_fwd.hpp/ordered_map.hpp and make CI fail on new ones (#5715 item 4c)

ci_single_binaries ran IWYU via CMake's CXX_INCLUDE_WHAT_YOU_USE launcher
property, which only printed "Warning: include-what-you-use reported
diagnostics" without failing the build: CMake's own __run_co_compile
wrapper does not propagate the launched tool's exit code, so even
`-Xiwyu --error` could never fail `cmake --build` this way. Verified
this empirically by injecting a deliberately-unused #include and
confirming the build still exited 0.

Fix the findings from the last recorded run (issue #5715 item 4, log
35829411620):
- ordered_map.hpp: add <new> (placement new) and
  nlohmann/detail/abi_macros.hpp; drop <memory> (std::allocator is
  still visible transitively via <vector>, confirmed by full local and
  containerized test suite runs).
- json_fwd.hpp: drop <memory> (same reasoning). Keep every forward
  declaration IWYU wanted removed (adl_serializer, basic_json,
  json_pointer, ordered_map): this file's only job is to forward-declare
  them for downstream users, so "nothing in this TU uses them" is
  expected, not a real finding. Mark each with `// IWYU pragma: keep`.
- json.hpp: add <cmath>, <cstdint>, <set>, <type_traits>,
  <unordered_map>, and the detail/abi_macros.hpp, detail/input/json_sax.hpp,
  detail/meta/detected.hpp, thirdparty/hedley/hedley.hpp includes IWYU
  says it needs. Do NOT remove adl_serializer.hpp,
  detail/conversions/from_json.hpp, detail/conversions/to_json.hpp,
  detail/macro_unscope.hpp, or ordered_map.hpp as IWYU suggests: nothing
  else in include/nlohmann includes adl_serializer.hpp or
  ordered_map.hpp, so basic_json<>'s own default template arguments
  (JSONSerializer = adl_serializer, and ordered_json = basic_json<ordered_map>)
  would lose their complete type; detail/macro_unscope.hpp is what
  undoes the JSON_* macros detail/macro_scope.hpp defines earlier in
  this same file, and removing it leaks those macros into every
  translation unit that includes <nlohmann/json.hpp>. Verified by
  actually removing them in a scratch test: the header still "compiles"
  stand-alone but ordered_json and every macro-using translation unit
  break. Marked each `// IWYU pragma: keep`.

Enforce it with `iwyu_tool` (ships with IWYU, e.g. as /usr/bin/iwyu_tool
on Debian/Ubuntu) instead of relying on the launcher property: it reads
compile_commands.json (now exported project-wide under JSON_CI) and
does return a real exit code for its own analysis, independent of
CMake's wrapper. ci_single_binaries now runs it over every
src_single/*.cpp with `-Xiwyu --error`, so a *new* finding fails CI.

json.hpp itself is excluded from that hard gate: even after every fix
above, IWYU's suggestion for one remaining symbol (a container
`swap, operator!=` used somewhere via a templated comparator) is not
deterministic — repeated, otherwise-identical containerized runs
reported <set>, then <unordered_map>, then <map> as "the" header to
add/remove for the exact same source. Gating a whole CI job on a
nondeterministic suggestion would make ci_single_binaries flaky rather
than informative, so json.hpp keeps the existing informational warning
(still shown during its normal compile) without failing the build on
it. Every other one of the ~50 single-header checks is included in the
hard gate.

#5715 item 4c. 4a (scan-build) and 4b (Infer) are separate commits.

Verified: full local ctest suite (129/129) and the ci_single_binaries
target itself both green in a containerized silkeh/clang:dev run
(matching the actual CI job) after this fix; a deliberately-reintroduced
unused #include in ordered_map.hpp was confirmed to fail
`cmake --build ... --target ci_single_binaries` (exit 2) with this
change, and to pass without it, on the same container/IWYU version CI
uses. `make check-amalgamation` is clean. Compiled with Clang and GCC
at -std=c++11/14/17/20 locally with no new warnings.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:49:24 +02:00
Niels Lohmann
c261578431 Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer

The doc block of write_number() ended up above the byte_swap() helpers
added in #5286, about 80 lines from the function. It was also a plain
comment that Doxygen skips, said "write a number to output input", and
left BON8 out of the big-endian formats. Move it back onto
write_number() as a /*! block and fix the text.

write_bson() documented "@pre j.type() == value_t::object", but it
throws type_error.317 for every other type, and to_bson() relies on
that. Document the exception instead.

Explain why the CBOR binary subtype is always written with a 0xD8..0xDB
head and never in the one-byte tag form: binary_reader with
cbor_tag_handler_t::store only keeps those heads as a subtype, so
switching to write_cbor_head() would break round trips for subtypes
0..23.

Also fix the grammar of the to_char_type comment. Comments only; no
change in behavior, API or ABI.

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the duplicated UBJSON/BJData integer marker ladders

write_number_with_ubjson_prefix() (unsigned and signed overloads) and
ubjson_prefix() (number_integer and number_unsigned cases) each picked
the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their
own independent if/else ladder, and the values beyond 64 bits were
handled by a second, tag-dispatched pair of ladders. An optimized
container announces the marker of its first element via ubjson_prefix()
and then writes every element through write_number_with_ubjson_prefix(),
so the two had to be kept in lockstep by hand across four call sites.

Replace all of that with one ubjson_integer_prefix() built on
value_in_range_of<T>, and one write_ubjson_integer_payload() that
writes the value (or, for 'H', the decimal digits) for a given marker.
write_number_with_ubjson_prefix() and ubjson_prefix() keep their
signatures and now just call these two helpers.

Behavior, the public API and the ABI are unchanged. Verified with a
new regression test covering scalars and $-optimized arrays/objects at
every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for
to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing
to_ubjson/to_bjdata output before and after over the json_test_data
corpus (bit-identical).

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove dead get_char parameters in binary_reader

The non-recursive rewrite of the binary readers (#5505, #5506, #5507)
left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char
parameters dead: parse_cbor_internal() has one caller and it always
passes true, and parse_ubjson_internal() has one caller and it always
uses the true default. Both parameters, and the @param docs describing
the "reuse the last character" mode they used to select, no longer
correspond to anything.

Drop both parameters, initialise fetch/prefix unconditionally, and
update the two call sites in sax_parse(). parse_cbor_value()'s and
get_ubjson_string()'s own get_char parameters are unrelated and are
left alone; both still have a false caller.

Also delete a stray `@return whether a valid MessagePack value was
passed to the SAX parser` doxygen block that sits directly above
parse_msgpack_value()'s real doc comment, a leftover of the same
rewrite.

Behavior, the public API and the ABI are unchanged; these are private
members of detail::binary_reader. Verified by compiling with
-Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and
unit-msgpack (offline, against the stubbed test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Share the IEEE half-precision decoder between CBOR and BJData

binary_reader had two ~45-line copies of the IEEE 754 half-precision
decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is
normalised, the two blocks were identical except for the byte order
used to assemble the 16-bit half (CBOR is big endian, BJData is little
endian). Any future change to half-float decoding had to be made and
kept in sync in both places.

Add one get_half_float(format, little_endian) helper that does the two
get()/unexpect_eof() reads, assembles the half in the requested byte
order, decodes it per RFC 8949 Appendix D, and calls sax->number_float.
Both cases now just call it with their byte order; the BJData case
keeps its bjdata-only guard.

Behavior, the public API and the ABI are unchanged. Verified with a
scratch probe comparing the old and new decoders bit-for-bit (NaN by
isnan()) over all 65536 wire byte pairs, in both formats, and by
running unit-cbor and unit-bjdata (offline, against the stubbed
test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the MessagePack unsigned-integer writer ladder

The number_integer (non-negative branch) and number_unsigned cases in
write_msgpack() each held their own copy of the fixint/uint8/16/32/64
ladder, kept in lockstep only by a comment ("we used the code from the
value_t::number_unsigned case here"). Both copies mixed union members:
the signed copy compared number_unsigned but wrote number_integer, and
vice versa.

Extract write_msgpack_unsigned(std::uint64_t), mirroring how
write_cbor_head() already avoids the same duplication for CBOR, and
call it from both cases. Each case now reads only its own active
union member. Output bytes are unchanged for the default 64-bit
number types.

#5710 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Unify float marker selection and fix the long double compile error

Four formats picked between a float32 and float64 marker through four
different helper styles: dummy-argument overloads for CBOR and
MessagePack, an std::is_same template for BON8, and a runtime if-chain
on input_format_t for write_compact_float(). With number_float_t set
to long double, to_cbor, to_msgpack and to_ubjson failed inside the
library with "call to 'get_cbor_float_prefix' is ambiguous", while
to_bson kept working because write_bson_double() takes a plain double.

Change write_compact_float() to take the two marker bytes directly
(each of its three callers already knows them at compile time) instead
of an input_format_t it only forwarded, and delete the now-unused
get_cbor_float_prefix(), get_msgpack_float_prefix(),
get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn
the two get_ubjson_float_prefix() overloads into one template. Both
write_compact_float() and get_ubjson_float_prefix() now report an
unsupported number_float_t with a static_assert naming the requirement,
rather than an ambiguous-overload error; the assert lives in the
function body, not the class scope, so to_bson with long double is
unaffected.

Verified with a probe basic_json<..., long double>: to_bson still
compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail
to compile with the new static_assert message.

This changes the text of an existing compile error for users with an
unsupported number_float_t (documented as a public-API-visible change
in #5710).

#5710 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the BJData ndarray writer's dtype dispatch and drop <map>

write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on
every call just to translate the _ArrayType_ name to a dtype marker
(the only reason binary_writer.hpp included <map>), then mapped dtype
to C++ type twice more: once as a switch for the range-check pass and
once as a separate if/else chain for the write pass, with nothing
checking that the two agreed. The caller also ran three at() lookups,
and the callee called value.at(key) about ten more times for the same
three members.

Replace the map with bjdata_ndarray_type_marker(), a plain string
comparison chain (a C++11 constexpr function cannot contain a switch,
so this mirrors binary_reader's own static table style). Replace the
switch/if-chain pair with one write_bjdata_ndarray_elements() that
switches on dtype once and calls a per-type helper -
write_bjdata_ndarray_element<T>() for the eight integer dtypes and
write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag
selecting the range check or the actual write, so the two passes can
no longer disagree on the type. _ArrayType_, _ArraySize_ and
_ArrayData_ are now looked up once into references, and the four
header marker bytes ('[', '$', '#') are written through to_char_type()
like the rest of the UBJSON/BJData writer.

The 'd' (single-precision) rule is left exactly as before, since #5707
is expected to change it separately.

Verified byte-for-byte identical output before/after for every dtype
(including the Draft 2/Draft 3 'byte' fallback and the use_count/
use_type combinations) via a standalone probe, plus round-tripping
through from_bjdata().

Overlaps #5707, which is expected to touch the 'd' dtype case, and
#5518, which is expected to move the write_bjdata_ndarray() call site.

#5710 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Assert that write_bson_document() consumes every calc_bson_sizes() entry

calc_bson_sizes() and write_bson_document() are a hand-synchronized
pair of passes over the same object/array tree, introduced by #5553:
the size pass appends to nested_sizes in visiting order, and the write
pass consumes the table by position with nested_sizes[next_size++].
Nothing checked that the write pass consumed the whole table. If a
future change touched only one of the two passes - for example to skip
or reject an entry - every later size prefix in the document would be
silently wrong.

Add JSON_ASSERT(next_size == nested_sizes.size()) where
write_bson_document() returns, so such a future drift between the two
passes is caught immediately (JSON_ASSERT expands to nothing in
release builds using assert(), and the fuzzers/tests already build
with it enabled). The two passes agree today, so this changes nothing
observable; it only guards against the risk described in #5710 item 5.

Extracting a shared stepper for the two passes (the second half of the
proposed change) is left for a follow-up: it only saves ~30 lines and
the issue asks for it only if the result reads clearly, which needs
more room to get right than a mechanical cleanup pass allows.

#5710 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the BJData lookup tables static functions instead of members

binary_reader held bjd_optimized_type_markers and bjd_types_map as
non-static const members (12 string_t objects for the type-name table),
built and destroyed on every from_cbor/from_msgpack/from_bson/
from_ubjson/from_bon8/from_bjdata call even though only from_bjdata
ever reads them. They also needed the #define/decltype/#undef
workaround from #3637 and two NOLINTNEXTLINE suppressions, and
binary_writer already carries the same two lists in another form
(is_bjdata_excluded_type_marker() and a local std::map in
write_bjdata_ndarray(), the latter removed by the item-4 commit), so
the excluded-marker lists could drift apart.

Replace bjd_optimized_type_markers with static constexpr
is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain
as binary_writer's is_bjdata_excluded_type_marker(). Replace
bjd_types_map with a non-constexpr static bjd_type_name(char_int_type)
switch returning nullptr for an unknown marker (a C++11 constexpr
function cannot contain a switch). Delete both
JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the
NOLINTNEXTLINE suppressions, detail::make_array() (no longer used
anywhere), and the now-unused <algorithm> and <array> includes.

Update the two call sites (the ND-array excluded-type check and the
_ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's
"LUT arrays are sorted" section, which only checked the two tables'
internal ordering, with a check of all 12 type names and all 8
excluded markers against both new functions.

#5711 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read CBOR's 1/2/4/8-byte argument through one helper

parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte
big-endian unsigned integer" ladder four times over:
- twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore
  branch and once, nearly identically, in the ::store branch (~90
  lines to read one integer);
- twice more for container lengths, once for array heads 0x98-0x9B and
  once for map heads 0xB8-0xBB, where the 1/2-byte forms called
  enter_array()/enter_object() directly and the 4/8-byte forms
  additionally went through get_cbor_container_size().

Add get_cbor_argument(std::uint64_t&), reading the width selected by
current & 0x1F via the same get_number() calls as before (so EOF is
reported exactly as before), and route all four sites through it:
- 0xD8-0xDB now read the argument once per branch instead of switching
  on `current` a second time; behavior split cleanly from embedded tags
  0xC0-0xD7 (tag value in the head, no argument to read), which is now
  its own case block that no longer has to fall into the ::store
  switch's "default" case to reach the same tag_pending = true; return
  true; outcome.
- 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always
  going through get_cbor_container_size() (harmless for 1/2-byte
  lengths, which already always fit).

Verified byte-for-byte identical behavior before/after with a
standalone probe covering embedded and multi-byte tags under all three
tag_handler_t settings, a tag over a byte string (subtype path),
truncated tag/length arguments of every width, and array/map lengths
of every width, including the out_of_range.408 "excessive size" case:
same exceptions, same messages, same chars_read, same successful
results.

Left the string/byte-string length ladders in get_cbor_string()/
get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is
expected to touch them separately.

Overlaps #5601 (adds a branch right above the embedded-tag case) and
#5607 (touches the integer cases 0x18-0x1B, which share this ladder's
shape in separate hunks).

#5711 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add leave_container() to match enter_container()

Every container is opened through enter_container(), whose docs
promise that a check placed there runs before every start event. The
close side had no equivalent: the same
"container_stack.pop_back(); dispatch to end_object() or end_array()"
sequence was written out separately in BSON, CBOR, MessagePack,
UBJSON/BJData and BON8, each copying the pattern of keeping an
is_object flag around the pop_back() that would otherwise invalidate
a reference to it. A check needed on close would have had to be added
in five places, and a sixth copy could go unnoticed.

Add leave_container() next to enter_container(), doing the same
pop-then-dispatch, and replace the five sites with it. Each site keeps
its own surrounding logic (BSON's check_bson_document_size() call
before popping, MessagePack's is_object copy used again below,
UBJSON/BJData's remaining-container handling after popping, BON8's
top used again below); only the repeated pop/dispatch line pair is
now shared.

Verified all six binary-format unit suites and unit-regression2's
deep-nesting tests (dependent count/reuse count and the bjdata ndarray
depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan.

Overlaps #5601, which is expected to add a sixth close site in its own
skip loop; that site can route through leave_container() too once it
lands.

#5711 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop passing the input format to sax_parse() when the reader already has it

binary_reader's constructor stores the format in the input_format
member, and sax_parse(format, sax_, strict, tag_handler) took the same
value again purely to dispatch on it. Every in-tree caller passed the
same value both times (all 16 from_cbor/from_msgpack/from_ubjson/
from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three
public basic_json::sax_parse() overloads), so nothing was broken
today, but a caller of the detail class directly (only reachable via
JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could
pass a mismatched pair - say bjdata to the constructor and ubjson to
sax_parse - and dispatch on one format while applying the other
format's rules; the default-constructed input_format_t::json reader
would additionally hit JSON_ASSERT(false) in exception_message() on
its first error.

Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the
existing overload with the stored input_format, and switch every
caller to it: the 16 from_*() sites (keeping their
`// cppcheck-suppress[accessMoved]` comments) and the three
basic_json::sax_parse() overloads, all of which already had the format
available from their own `format` parameter. The four-argument overload
is kept for anyone still calling it, now with
JSON_ASSERT(format == input_format) so a mismatch fails immediately
in a debug build (assert-enabled binaries, including the fuzzers and
test suite) instead of misbehaving; verified with a probe that
constructs a reader for one format and calls the explicit overload
with another, which aborts on that assertion as expected.

Removing or asserting against the constructor's input_format_t::json
default, which would affect direct detail users, is left as a separate
decision per #5711 item 5.

Overlaps #5601, which is expected to add an AllowRecovery template
parameter to sax_parse() and touch these same call sites in json.hpp.

#5711 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks

get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently
sized signed integer and then repeated the same "reject negative with
error 113" check; only 'L' additionally checked value_in_range_of for
the out_of_range.408 case. Any change to that error path had to be
made four times.

Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read,
the negative check and the range check once, and route all four
markers through it. The range check is a no-op for 'i'/'I'/'l' (their
values always fit std::size_t) and only live for 'L' on a 32-bit
std::size_t target, matching today's behavior exactly.

In the ndarray dimension-product loop, the preceding loop already
returns early on any zero dimension and result starts at 1, so `i > 0`
in the pre-multiplication overflow check was always true, and
`result == 0` in the post-multiplication check could not be reached
either: two positive factors whose product does not overflow (as the
pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&`
and narrow the post-check to `result == npos`, the one case the
pre-check cannot rule out (an exact, non-overflowing match with the
sentinel reserved for unknown-size containers), with a comment
explaining why.

Verified byte-for-byte identical behavior before/after with a
standalone probe covering negative counts for every marker, a matching
positive count, and ndarray inputs, plus the full unit-ubjson and
unit-bjdata suites (same assertion counts as before this change).

Overlaps #5601 (rewrites the four parse_error calls and the overflow
checks touched here) and #5607/#5707 (touch neighboring lines in the
same functions).

#5711 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop redundant format parameter and dummy float argument (review)

binary_reader::sax_parse(format, ...) only ever had to equal the format
given to the constructor, which it asserted. With every caller already
on the format-less overload, remove the four-argument overload and
dispatch on the stored input_format directly. binary_reader is a
detail class, so this is not a public API change.

get_ubjson_float_prefix() took a value only to deduce its type; make
the type an explicit template argument instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:48:32 +02:00
Niels Lohmann
35802e78d6 Remove dead Makefile targets; document macro_builder; tidy serve_header (#5735)
* Remove dead doctest help entry and pretty_format target from Makefile

The top-level Makefile still carried three leftovers:

- The help text listed a "doctest" target that was removed in #4560,
  so "make doctest" fails with "No rule to make target". The example
  check now runs as "make check_output -C docs".
- "pretty_format" ran clang-format on all sources, but .clang-format
  was deleted in #4573, so the target reformatted everything in the
  default LLVM style, against the Artistic Style formatting that
  "make pretty" applies and CI enforces.
- "clean" removed benchmarks/files/numbers/*.json, a directory that no
  longer exists since the benchmarks moved to tests/benchmarks (#3462).

Only maintainer tooling changes; the library is not affected.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document tools/macro_builder and tidy up serve_header.py

tools/macro_builder generates the NLOHMANN_JSON_EXPAND,
NLOHMANN_JSON_GET_MACRO and NLOHMANN_JSON_PASTE* macros in
macro_scope.hpp, but nothing referred to it. Add a README that explains
what it generates, how to run it and where the output goes, and which
dependent tables (NLOHMANN_JSON_DOUBLE_PASTE, NLOHMANN_JSON_TYPE_BODY)
are maintained by hand. Point to it from a comment above
NLOHMANN_JSON_EXPAND. The generator itself is unchanged; following the
README reproduces the header byte for byte.

In serve_header.py, drop the LGTM suppression (LGTM.com shut down in
2022), replace the """.""" placeholder docstrings with real ones, and
import socket and ssl at module level. DualStackServer.server_bind uses
socket, which was only imported under __main__; when the module was
imported instead, the NameError was swallowed and IPV6_V6ONLY was not
cleared.

The header change is a comment only; behavior, API and ABI are
unchanged.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Hash every release_files artifact, not a hardcoded subset

The `release` target signed and copied json_fwd.hpp into release_files
alongside json.hpp, but the shasum line that writes hashes.txt only
listed json.hpp, include.zip and json.tar.xz. Users could not verify
the published json_fwd.hpp against hashes.txt.

Hash every file in release_files except the .asc signatures instead
of naming files by hand, so a newly shipped header (such as the
json_literals.hpp that #5610 adds to this target) cannot be missed
again.

Only affects the generated hashes.txt release artifact; the library
itself is unaffected.

Overlaps #5610, which touches the same lines to add json_literals.hpp
to the release target.

#5717 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the broken fuzz_testing* Makefile targets

fuzz_testing and fuzz_testing_{bon8,bson,cbor,msgpack,ubjson} seeded
fuzz-testing/testcases from tests/data, which was removed in dbf1a1f41
(2020) when the test data moved to the external json_test_data repo.
The find command found nothing, but the pipeline's exit status was
that of xargs, so the recipe still reported success with an empty
corpus, and the printed afl-fuzz command would refuse to start.

The recipes were also six near-identical copies with unquoted -name
patterns, used the legacy CXX=afl-clang++, and fuzzing-start/stop were
missing from both the help output and .PHONY. tests/fuzzing.md already
documents the working flow (download json_test_data, then
`make -C tests fuzzers`), so replace the six broken targets and their
help lines with a single pointer to that document instead of trying
to keep six copies of a fragile shell pipeline in sync.

This does not affect OSS-Fuzz, which builds through tests/Makefile.

Overlaps #5621, which adds a seventh copy of the same broken line for
fuzz_testing_json_view.

#5717 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove stale Travis comment above check-amalgamation

check-amalgamation carried "Note: this target is called by Travis",
left over from before the project switched off Travis CI. The prior
Makefile cleanup commit removed the other stale Travis-era leftovers
(the doctest help entry, pretty_format, and the benchmarks/ path in
clean) but missed this comment.

#5717 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale install/usage instructions in the vendored amalgamate README

tools/amalgamate/README.md is the unmodified upstream text and no
longer matches how the tool is used here:

- It named a Bitbucket origin that no longer exists; CHANGES.md
  already tracks the GitHub mirror commit this copy is based on.
- It asked for Python 2.7, but CI and the Makefile run the script
  with python3.
- It told readers to run ./test.sh (not vendored) and install to
  /usr/local/bin; in this repository the tool runs through
  `make amalgamate`.
- Its usage synopsis showed `-v` taking no argument, but the script's
  own argparser requires `choices=["yes", "no"]`, so that form fails
  with "argument -v/--verbose: expected one argument". The Makefile
  calls it as `--verbose=yes`.
- It pointed at test/source.c.json and test/include.h.json, which are
  not vendored; the configs actually used are config_json.json and
  config_json_fwd.json.

Rewrote only the Installing and Using sections to match; left the
"Here be dragons" caveats and the rest of the vendored code untouched
to avoid diverging further from upstream.

Overlaps #5615, which edits amalgamate.py, this README and CHANGES.md.

#5717 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Derive generate_natvis.py's ABI tag list and version from abi_macros.hpp

generate_natvis.py hard-coded abi_tags = ['_diag', '_ldvcmp', '_dp',
'_bics', '_psp', '_snul'] and required --version on the command line.
The source of truth is include/nlohmann/detail/abi_macros.hpp: the
NLOHMANN_JSON_ABI_TAG_* defines, the argument order of
NLOHMANN_JSON_ABI_TAGS_CONCAT, and NLOHMANN_JSON_VERSION_MAJOR/MINOR/
PATCH. Nothing checked that the copies stayed in sync, and they have
drifted apart before: _dp was added in #4517 but missed here until
#5544, and 3.11.3 shipped json_abi_v3_11_2 namespaces (#4340).

Parse the tag list (in NLOHMANN_JSON_ABI_TAGS_CONCAT order) and the
version from abi_macros.hpp instead of hard-coding them. Make
--version optional (falling back to the parsed version) and default
the output directory to the repository root the script lives in.

Add a "natvis" Makefile target that runs the script, and extend
check-amalgamation to regenerate nlohmann_json.natvis and fail on a
diff, the same way it already does for the amalgamated headers and
BUILD.bazel. Wire the same regeneration into check_amalgamation.yml,
using the tool copy checked out from develop (as the workflow already
does for amalgamate.py) and installing jinja2 from
tools/generate_natvis/requirements.txt. Update the tool's README to
say it must be re-run after adding an ABI tag or bumping the version.

Verified: a run against develop produces no diff (with either the
default or an explicit --version 3.12.0); adding a dummy
NLOHMANN_JSON_ABI_TAG_* without a matching #define makes the script
fail loudly instead of silently omitting the tag; xmllint --noout
passes on the regenerated file; and running the script from a
directory other than the one being checked (simulating the workflow's
separate tool checkout) against this repository root also produces no
diff.

Overlaps #5600, which added _ekmo to the same hand-written abi_tags
line and regenerated the file.

#5717 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Strip the leading "./" find(1) prefix from release hashes.txt entries

bc6e7db72 (#5717 item 1) switched the release target's shasum line from
naming files by hand to $$(find . -type f -not -name '*.asc' | sort),
so a newly shipped header is hashed automatically. Run from inside
release_files, that find prints paths as "./json.hpp" instead of
"json.hpp", so hashes.txt lists "./json.hpp" etc. instead of the plain
filenames it always used. shasum -c still verifies "./json.hpp" fine,
but it is a needless cosmetic regression for anyone reading the file
or matching it against release notes.

Strip the "./" prefix with sed before sorting, keeping the filenames
exactly as before while still hashing every artifact automatically.

Review fix for #5717 item 1 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix check_amalgamation.yml: pass --version to generate_natvis.py

bff45f111 (#5717 item 3) made --version optional in
tools/generate_natvis/generate_natvis.py and wired the workflow's new
"Regenerate nlohmann_json.natvis" step to call it without --version,
relying on the script deriving the version from abi_macros.hpp itself.

But NATVIS_TOOL_DIR is checked out from develop, the same way TOOL_DIR
already is for amalgamate.py, precisely so an in-flight PR's tooling
changes cannot mark themselves clean. Until this PR (or an equivalent)
merges to develop, that checkout is the old generate_natvis.py, whose
--version argument is still required=True. The new step's invocation
of "generate_natvis.py $MAIN_DIR" (no --version) then fails argparse
on this PR's own CI run with "the following arguments are required:
--version", before the check ever gets to compare output.

Extract the version from $MAIN_DIR's own abi_macros.hpp in the
workflow and always pass it as --version. That satisfies the old
script's required argument and is accepted as an explicit override by
the new one, so the step behaves the same whether NATVIS_TOOL_DIR holds
the pre- or post-merge tool, and stays correct for later PRs that bump
the version.

Verified by running the workflow step's shell logic locally against
both the pre-#5717 generate_natvis.py (checked out at 633de8e44) and
the new one: both produce the identical nlohmann_json.natvis as the
committed file.

Review fix for #5717 item 3 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extend tools/macro_builder to also generate DOUBLE_PASTE and TYPE_BODY

tools/macro_builder only emitted NLOHMANN_JSON_EXPAND..PASTE64, so two
other tables that scale with the same max_args stayed hand-maintained
with nothing checking them: NLOHMANN_JSON_DOUBLE_PASTE (added by hand
in #4563 for the *_WITH_NAMES macros) and the 64-slot dispatch table of
NLOHMANN_JSON_TYPE_BODY (the #4041 zero-member/one-or-more-member
switch). Both tables pass one macro name per slot to the same
NLOHMANN_JSON_GET_MACRO dispatch as PASTE, so they can drift out of
sync with max_args exactly the way _dp did in the ABI tag list fixed
by #5544.

Extend main.cpp with build_double_paste_code() (same recursive-doubling
shape as build_paste_code(), but DOUBLE_PASTE consumes two arguments
per member, so an even slot index falls back to the next lower odd
DOUBLE_PASTE<N>) and build_type_body_table() (max_args - 1 MEMBERS
slots and one trailing EMPTY slot, 8 per line, matching how it is
written by hand today). Add a "type_body" argument that selects the
TYPE_BODY block, since it lives at a separate location in
macro_scope.hpp from the EXPAND..DOUBLE_PASTE63 block; plain invocation
is unchanged apart from covering the extended range. No longer emit
the tool's old trailing blank line, so its output is directly diffable
without post-processing.

Verified with c++ -std=c++11: running the tool (with and without
"type_body") and piping the raw output through the pinned astyle
reproduces both blocks of the current macro_scope.hpp byte for byte.
tests/src/unit-udt_macro.cpp (all NLOHMANN_DEFINE_TYPE_*/_WITH_NAMES/
zero-member variants) passes unchanged under -std=c++11 and -std=c++17
with -fsanitize=address,undefined.

Add a "macro_builder_check" Makefile target that builds main.cpp,
regenerates both blocks into a scratch directory inside the repository
(astyle's --project lookup needs the target files under the same tree
as .astylerc, unlike an external /tmp directory), and diffs them
against the corresponding ranges of macro_scope.hpp; wire it into
check-amalgamation next to the natvis check. Wire the same regeneration
into check_amalgamation.yml, splicing the (still unindented) generated
blocks back into the PR's own macro_scope.hpp before the existing
astyle/amalgamation step runs, so that step's own tree-wide astyle
pass both indents them and folds any drift into the amalgamation
patch/diff the workflow already produces.

Unlike amalgamate.py and generate_natvis.py, this step builds
tools/macro_builder/main.cpp from the pull request's own checkout
($MAIN_DIR) rather than a separate checkout of tools/ at develop: this
tool has no independent source of truth to regenerate against (its
README documents that it must reproduce macro_scope.hpp byte for
byte), so a develop-pinned copy would only reproduce the
generate_natvis.py trap fixed in a previous commit on this branch,
where a PR that teaches the tool to cover more of the file fails its
own CI until that PR merges and updates the develop copy.

Add tools/macro_builder/README.md documentation for both new tables
and the two-invocation usage, and a short pointer comment above
NLOHMANN_JSON_TYPE_BODY (the EXPAND pointer already covered the first
block; extended its wording to include DOUBLE_PASTE63).

Closes #5717 item 5 in full, completing what the documentation-only
"Document tools/macro_builder..." commit already on this branch left
open (that commit's README/pointer-comment half stands; it also covers
item 7).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:21:27 +02:00
Niels Lohmann
7d7055ec50 Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations

Constructing a basic_json from another specialization (json to
ordered_json or back, also via get<ordered_json>()) converted every
container with its range constructor, which calls the converting
constructor for each element. The call stack therefore grew with every
nesting level, and a value nested some 30,000 levels deep overflowed it.

The conversion now bounds its descent the way the copy constructor does
since #5387: the first 128 levels are converted exactly as before, and
below that convert_iteratively() finishes the value with an explicit
stack. It builds each container bottom-up from its converted elements
with the container's range constructor, so member order and keys that
become equal are handled as before, and it gives a value its type only
once its container exists, so an exception leaves nothing behind that
cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions
(JSON_DIAGNOSTIC_POSITIONS) are set for every value.

Converting a null value no longer resets its positions: the constructor
assigned null to a value that already was null, which swapped in the
positions of the temporary.

Fixes #5650.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Explain why converting null keeps positions and why next is a reference

Review feedback on #5723 (gregmarr): clarify in comments that the
converting constructor has already copied the positions of val, which
the null case keeps like every other case, and that next must be a
reference into pending so that ++next advances the stored iterator.

Comments only; no code change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Refer to recursion_depth_limit() in the convert_structured() docs

The comment still named nesting_depth_limit, which #5637 removed on
develop in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Advance the pending iterator through pending.back() and shorten the null comment

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:36:04 +02:00
Niels Lohmann
dcc81f43f0 Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
  number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names

The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.

Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:09 +02:00
Niels Lohmann
383ce0b040 Cover the edit storage in tests
Test edits of an empty document and assignments through a view of a
value that is no longer part of the document; copy the entries of a
block through deref() (one path for links and values); mark the
4 GiB limit and the returns of find_parent() that no document reaches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:06 +02:00
Niels Lohmann
51ee239b3c Address the clang-tidy findings of the editable documents
Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:04 +02:00
Niels Lohmann
6e50766dc3 Add editable documents with set() and push_back()
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:

- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
  object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
  the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)

Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.

Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.

Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:03 +02:00
Niels Lohmann
d4fab0971f Walk the index of json_view through a navigation policy
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:

- node flags `edited` (a string or number token in the edit arena),
  `moved` (the elements of an array/object live in a separate sequence),
  and `is_new` (no source position), and link nodes (kind_link) that
  stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
  values, and the edit arena

materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:02 +02:00
Niels Lohmann
eede67ca92 Fix old clang: do not declare the defaulted document_data() noexcept
With the nested struct object_index, clang 4 (and, by the same bug, the
clang 3.x of ci_test_compilers_clang) rejects the explicitly noexcept
defaulted constructor: "default member initializer for 'indexes' needed
within definition of enclosing class 'document_data' outside of member
functions". Nothing depends on the constructor being noexcept, so let it
take the implicit exception specification.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:01 +02:00
Niels Lohmann
82b31f31b8 Fix CI: useless casts of the key hash of the view's object index
GCC -Werror=useless-cast on Linux x86-64 rejects
static_cast<std::size_t>(key_hash(...)): the call returns a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Store the hash in a variable
and cast that, which GCC does not report.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:00 +02:00
Niels Lohmann
904c8c6710 Index the large objects of json_view
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.

Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.

Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:58 +02:00
Niels Lohmann
1f0c3be6f3 Address the clang-tidy findings of the SIMD scan
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:56 +02:00
Niels Lohmann
c3b51addbf Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:55 +02:00
Niels Lohmann
ecf9df7c45 Convert the floats of json_view from the digit layout
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For floats and doubles with at most 19
digits, the value is now read from that layout: the digits eight at a time,
without scanning the token, and rounded by the library's conversion core
(detail::decimal_to_float(): Clinger's fast path where both operands are
exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19
digits). It rounds correctly, so the values are those of parse(); other
tokens and types keep the library's conversion of the whole token.

get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s.

Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and
those of float) to the bit-for-bit comparison with parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:51 +02:00
Niels Lohmann
a7fa8d04e8 Address the cpplint findings of json_view's comparisons
compare.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:49 +02:00
Niels Lohmann
677507137b Address the clang-tidy findings of the comparisons
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:49 +02:00
Niels Lohmann
e842f3a68f Add comparisons to json_view
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.

Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:47 +02:00
Niels Lohmann
bfb2b0cb48 Mark the cases of the view's serializer that tests cannot reach
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:01:45 +02:00