mirror of
https://github.com/nlohmann/json.git
synced 2026-10-01 22:45:17 +00:00
89bf08760f5c7c6740e9b3e50d64e52c40970568
291 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5277335a9e |
Merge branch 'json-view/12-view-values' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
05b6cd0892 |
Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
95d10dab70 |
Merge branch 'json-view/10-view-document' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
371a8a3d9f |
Merge branch 'json-view/08-view-builder' into json-view/10-view-document
Signed-off-by: Niels Lohmann <mail@nlohmann.me> # Conflicts: # Makefile # cmake/ci.cmake |
||
|
|
9d44e3f359 |
Merge branch 'develop' into json-view/02b-float-parser
Conflicted only in tests/src/unit-class_lexer.cpp, where develop's #5737 lint fix (CAPTURE(x); -> CAPTURE(x)) collided with this PR's rewrite of the Eisel-Lemire float tests; kept the PR's new tests and applied the lint-fixed CAPTURE style. single_include regenerated via make amalgamate. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9d88ead578 |
Clarify that the strtold fallback substitutes the locale's decimal point
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
a32f61eb98 | Fix update() and merge_patch() when the argument is *this or one of its members (#5678) | ||
|
|
1d675cdb46 |
Fix CI jobs that check less than they claim; move arm64 to GitHub (#5733)
* Fix the ci_cmake_flags wiring so every option is checked
The CMake 3.31.6 flag list referred to itself before it was defined,
so only JSON_BuildTests was checked with that version. The targets for
the CMake running the build ("_2") were created but never added to
ci_cmake_flags, and the three versions shared one build directory.
JSON_StrictNulHandling was not in the list at all.
Use the 3.5.0 list for 3.31.6, add JSON_StrictNulHandling, and create
one ci_cmake_flag_<flag> target per option for the running CMake with
its own build directory. Also use the function parameter in the
COMMENT, refresh the stale version comment, and let ci_clean remove
the downloaded cmake-<version> directories instead of the long-gone
cmake-3.5.0-Darwin64.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix ci_test_clang_libcxx_cxx* jobs silently building without warnings
CMake only seeds CMAKE_CXX_FLAGS from the CXXFLAGS environment variable
when the cache entry is unset, so the explicit -DCMAKE_CXX_FLAGS="-stdlib=libc++"
argument made it ignore CXXFLAGS="${CLANG_CXXFLAGS}" entirely. The six
ci_test_standards_clang (..., libcxx) jobs therefore compiled without
-Weverything/-Werror while their libstdc++ siblings did use them.
Pass -stdlib=libc++ through the same CXXFLAGS value instead of a separate
-D argument, and give the target its own build directory
(build_clang_libcxx_cxx${CXX_STANDARD}) so it no longer shares a CMake
cache with the libstdc++ variant. Suppress the resulting
-Wthread-safety-negative finding from libc++'s std::mutex annotations,
which fires on doctest's reporters in this translation unit only.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the no-op AppVeyor with_win_header job
The with_win_header matrix entry patched Windows.h into
single_include/nlohmann/json.hpp before building, but JSON_MultipleHeaders
has defaulted to ON since #3532 (2022-06), so CMakeLists.txt points the
tests at include/ and the patched single header is never compiled. The
job has been a no-op VS2015 build since then.
Windows.h coverage already exists through tests/src/unit-windows_h.cpp
(#3631), which runs in every MSVC job. Delete the dead matrix entry and
its before_build steps, and cite unit-windows_h.cpp from the QA page.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove unused ci_oclint and ci_pvs_studio targets
No workflow invokes ci_oclint, ci_pvs_studio, or their tool discovery.
ci_oclint also had a side effect on every JSON_CI configure: it copied
the single header into src_single/all.cpp and added an add_executable()
for it without EXCLUDE_FROM_ALL, so a plain build compiled a 1.2 MB
translation unit that only that unused target consumed. ci_pvs_studio
duplicates the Makefile's pvs_studio target, which is kept.
Also drop the duplicate --check-level=exhaustive flag passed twice to
the same ci_cppcheck invocation.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop Dependabot from proposing astyle bumps
astyle is deliberately pinned at 3.4.13 because newer versions reformat
unrelated lines and this version defines the formatting that
check_amalgamation.yml enforces. Without an ignore rule, Dependabot
keeps opening PRs for every new astyle release (most recently #4580,
#4942, #5445, #5448), each of which fails the amalgamation check and
gets closed unmerged.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move Linux arm64 CI from dead Cirrus CI to ubuntu-24.04-arm
Cirrus CI stopped reporting check runs on develop sometime after
|
||
|
|
791cd88dfc |
Fix raw-TeX formulas and other documentation infrastructure debt (#5736)
* Fix math formulas rendering as raw TeX in the published docs The privacy plugin self-hosts MathJax 2.7.0 but drops its ?config=TeX-MML-AM_CHTML query string, so the rehosted script loads no input jax and the 10 formulas across 6 pages render as raw TeX to readers. Remove pymdownx.arithmatex and the MathJax extra_javascript entry, and rewrite the formulas in plain HTML (<sup>, <i>) instead. This also drops a nine-year-old third-party script from every page. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix publish_documentation triggers and persist unneeded git credentials publish_documentation.yml only triggered on docs/mkdocs/** pushes, but the site also embeds .github/CODE_OF_CONDUCT.md, CONTRIBUTING.md, SECURITY.md, cmake/{clang,gcc}_flags.cmake, .clang-tidy, tools/astyle/.astylerc and tests/fmt_formatter/project/main.cpp via pymdownx.snippets, so changes to those files never republished the site. Extend the path filter to cover them, and switch runs-on from the long-pinned ubuntu-22.04 to ubuntu-latest to match ci_test_documentation. Also add persist-credentials: false to the checkouts in ci_icpx and ci_nvhpc (ubuntu.yml) and msvc-vs2026/msvc-arm64 (windows.yml), none of which pushes with git, matching every other checkout in these workflows. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix example build: broken debug echo, deprecations hidden for everything docs/Makefile's debug echo used a space instead of a comma in $(call cxx_standard ...), so it always printed an empty standard. Every example was also compiled with -Wno-deprecated-declarations, which would silently hide an accidental deprecated-API call in any of them. Factor the duplicated compile flags into EXAMPLE_CPPFLAGS/ EXAMPLE_WARNFLAGS, build with -Werror=deprecated-declarations by default, and only allow the three examples that intentionally document deprecated API (the DEPRECATED_EXAMPLES list) to suppress it. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix check_structure.py NOLINT parsing and an off-by-one report line `line.strip("<!-- NOLINT")` followed by `.strip(" -->")` strips any of the characters in those sets from both ends, not a literal prefix; it only happened to work for "Examples". A NOLINT'd section name starting with N, O, L, I or T (e.g. "Notes", "Template parameters", "Iterator invalidation", "Literals") was silently mangled, so the suppression did not apply and the checker could report a spurious missing/misordered section. Parse the comment with a regex instead. The same fragile strip() pattern was used for heading text; replace it with a plain prefix slice. Also fix the admonition_title report, which used the 0-based line index while every other report in the file uses lineno+1. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix docset README's non-existent make target and stale fallback URL The README told readers to run `make nlohmann_json.docset`, but the Makefile's targets are `all`, `JSON_for_Modern_C++.docset` and `install_docset_zeal`; the documented command has failed with "No rule to make target" since #2967 (2021). Point the README at the real target and folder name. Info.plist's DashDocSetFallbackURL also still pointed at the old nlohmann.github.io/json/ URL instead of the canonical https://json.nlohmann.me/ from mkdocs.yml's site_url. Leave list_missing_pages/list_removed_paths alone: they may become redundant once #5638's check_docset() lands, which is a follow-up. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove two leftover Doxygen-era .link files from the examples directory parse__iterator_pair.link and parse__pointers.link each held only a Wandbox "online" permalink from the old Doxygen docs. #3071 deleted every other .link file in 2021; these two came in through a parallel PR (#3100) and were never referenced by any page, script or config. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix MacPorts CMake example to include its own snippet files The MacPorts "Example: CMake" block included integration/homebrew/example.cpp and integration/homebrew/CMakeLists.txt instead of the MacPorts files right next to it, a copy-paste slip from the Homebrew section. Nothing referenced integration/macports/CMakeLists.txt as a result. The page rendered correctly only because the homebrew, macports and vcpkg/CMakeLists.txt snippets are byte-identical, so a future edit to the MacPorts files would not have shown up on the page. Part of #5718 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Do not persist git credentials in publish_documentation's checkout The checkout step in publish_documentation.yml left the default persist-credentials: true, so GITHUB_TOKEN stayed writable in .git/config for the rest of the job (zizmor's artipacked finding). The Deploy documentation step authenticates through its own github_token input to peaceiris/actions-gh-pages and does not push with the checked-out credentials, so persist-credentials: false is safe here, matching every other checkout in the workflow set. Overlaps #5638, which edits this same checkout step (adds fetch-depth: 0); expect a rebase conflict there. Part of #5718 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Replace list_missing_pages/list_removed_paths with a comm(1)-based diff The docset Makefile's list_missing_pages ran one sqlite3 query per mkdocs page, and list_removed_paths nested a loop over all mkdocs pages inside a loop over all docset index paths (O(n*m) shell iteration). Issue #5718 item 5 suggested removing or reducing these targets once #5638's check_docset() lands, but that PR is still open and covers only API pages and macros, not the full page set these targets check. Replace the loops with two sorted path lists (DOCSET_PAGE_PATHS from mkdocs' markdown sources, DOCSET_INDEX_PATHS from the built docset index) compared with a single comm(1) call each, verified to produce output identical to the old loops against the current docSet.dsidx. The sed expression used '#' as its delimiter, which GNU Make reads as a comment character even inside a variable assignment, truncating the line and orphaning the closing paren of $(shell ...) ("unterminated call to function 'shell': missing ')'"). Use '@' as the delimiter instead. Part of #5718 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bc56ac6d9a |
Format the dump() example with the pinned astyle
The "check" job runs astyle over the documentation examples once it gets past the amalgamation step. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
c7e534d41e |
Document dump() of json_view
- API pages for dump, number_format, and operator<< of basic_json_view, linked both ways with the basic_json pages - the feature page describes document order and number_format::source - the examples show when the view helps: forwarding part of a message and writing numbers exactly as they were read Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
43e2d9c525 |
Document values and JSON pointers of json_view
- API pages for get, get_to, get_string, number_token, and value of basic_json_view; JSON pointer overloads of operator[], at, and contains; links both ways with the basic_json pages - the feature page describes which conversions copy nothing - the examples show when the view helps: strings without copies, numbers exactly as written, and paths into a large text Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e17e23a2b1 |
Fix CI: json_view access tests without exceptions and on clang 3.6
- ci_test_noexceptions: the element access tests compare the exceptions of json_view and basic_json through exception_of(), which catches them outside a CHECK_THROWS; with JSON_NOEXCEPTION the first one aborted the test. Compile those comparisons only with exceptions. - clang 3.6: value-initialize a const json_view, as in the tests of json-view/10-view-document. - Format three new documentation examples with the pinned astyle, which the "check" job runs once it gets past the amalgamation step. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
08dcdaf9b8 |
Document element access, lookup, and iteration of json_view
- API pages for operator[], at, front, back, find, contains, count, begin, end, cbegin, cend, items, and type_name of basic_json_view, linked both ways with the basic_json pages - the feature page and size() describe document order and duplicate keys - the examples show when the view helps: reading a few fields of a large text, probing optional members, and members in source order Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d9a71e7bc5 |
Document the node index of json_view on the architecture page
A new section describes the 16-byte node: its fields, how integers, floats, and object members are stored, how views navigate without pointers, and a worked example. The feature page and the pages of basic_json_document and node_count link to it where they mention the 16 bytes. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
061c30310e |
Document json_document and json_view
- API pages for basic_json_document and basic_json_view, one per member, and for the four aliases, each with an example - features/json_view.md: the problem the view solves, ownership and lifetime, what matches basic_json::parse() and what differs, and when to choose json, ordered_json, SAX, or the view - the examples show why one would use the view, not only how: borrowed vs. owned input, reading a few fields and materializing one subtree, reusing a document across many messages - registered in the mkdocs navigation, llms.txt, the docset, the exceptions page (out_of_range.416), architecture.md, the integration page, and the README; the yyjson credit is added to the README and license.md Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
ad73c78923 |
Export json_document and json_view from the C++20 module
The nlohmann.json module includes json_view.hpp in its global module fragment and exports basic_json_document, basic_json_view, and the four aliases next to basic_json, json, and ordered_json. features/modules.md lists them, and tests/module_cpp20 parses a document, so that a missing export fails the ci_module_cpp20 job. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
3b6ae43c53 |
Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions basic_json can be constructed from std::tuple<json&>, which it turns into a one-element array. Because of this, std::tuple picks its converting constructor that converts the whole source tuple instead of the element-wise one. As a result, std::tuple<const json&> built from std::forward_as_tuple(j) binds to a temporary (a compile error with libc++, a dangling reference with other standard libraries), and std::tuple<json> built the same way holds [j] instead of a copy of j. The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option JSON_DisableTupleReferenceConversion) removes the conversion from a one-element tuple holding a reference to the same basic_json type, so std::tuple converts element-wise. It is off by default, so existing behavior is unchanged. It does not change any function body and therefore is not part of the ABI tag. Fixes #2226 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert one-element tuples to arrays on every compiler to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }. With a single element that is itself a basic_json, Apple clang 15 and 16 treat j = {x} as a copy of x, so std::tuple<json>{true} became true instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new checks in unit-disable-tuple-reference-conversion and unit-regression2. The one-element overload that already handles JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a [string, value] element) explicitly, the same way the initializer-list constructor does. Use it unconditionally. The output is unchanged on compilers that already wrapped the element. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip json reference tuple tests on clang < 4 and GCC < 5 ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could not compile the new tuple tests. Creating a std::tuple of basic_json references, e.g. std::forward_as_tuple(j), makes these compilers instantiate basic_json's conversion operator for libstdc++'s internal tuple bases, which fails hard. This happens with and without JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these compilers, not of the new option. Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail, clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the checks that create such tuples; the is_constructible checks still run. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9c71689715 |
Convert long doubles under a multi-byte decimal point completely
The strtold fallback, which is left only for long double formats that are not binary64 (x87, binary128), substituted the first byte of the locale's decimal point for '.'. Under a locale whose decimal point is longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B), strtold stopped there and the value was truncated at the decimal point. A longer decimal point is now put into a copy of the token. The test "locale with a multi-byte decimal point" now compares the long double values with those of the "C" locale; with x87 long doubles it failed before. Fixes #5660. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
44ec53c77b |
Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple arm64) are now converted by the library itself, correctly rounded and independent of the locale and of the C and C++ libraries: - The token is split into sign, significand w (at most 19 digits), and decimal exponent q, using the positions of the decimal point and the exponent that the scanners already recorded, so no character is classified again. - Clinger's fast path where w and 10^|q| are exact. - Eisel-Lemire otherwise, now templated for binary32 and binary64. - For tokens with more than 19 digits whose w and w + 1 round differently, an exact big-integer comparison with the midpoint between the two candidates (the digit comparison of fast_float, simplified). This replaces the separate token walks of Clinger's fast path and of Eisel-Lemire, the significant-digit gate that avoided the former, and, for float and double, std::from_chars and the locale-aware strtod. std::from_chars and strtold remain only for other long double formats (x87, binary128, double-double) and for types that are not IEEE-754. Values are bit-identical to before wherever the previous conversion was correctly rounded; tokens converted in a locale with a multi-byte decimal point are now also exact. Overflow still gives out_of_range.406, underflow a signed zero. convert_float() is the entry point for other parsers of JSON text: it converts like the lexer, without allocation for binary32/binary64. Tests: exact-bit tests for double and float (ties, subnormal and overflow boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for binary32, the round trips of 200,000 doubles and 100,000 floats without declines, 508 generated hard cases with the expected bits of both formats (float_hard_cases.hpp) through the converter and both scanners, and JSON-level overflow/underflow checks for double and float. The locale tests now check the values in a locale with a multi-byte decimal point. Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float credit now names the digit comparison. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6a073dbae4 |
Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse error left the target holding whatever was parsed before the error instead of its previous value. With JSON_DIAGNOSTICS=1, that partial value also violated the class invariant, because the parent pointers of an array or object's elements are only set when the container is closed, which a failed parse never reaches; copying such a value then aborted in assert_invariant(). Fix it the way basic_json::parse() already handles this: parse into a temporary and move it into the target only once parsing succeeds, so the target is left unchanged if an exception is thrown. Fixes #5652. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
fdcc569eee |
Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.
Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.
Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).
Fixes #5666.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
6fc0d501f3 |
Move the user-defined string literals to <nlohmann/json_literals.hpp> and add JSON_NO_AUTOMATIC_UDLS (#5610)
* Add JSON_NO_UDLS to leave out the user-defined string literals The bodies of operator""_json and operator""_json_pointer call the parser, so every translation unit including the library instantiates it, even if it never parses anything. Defining JSON_NO_UDLS leaves the literals out entirely, which saves 15-35% compile time for such translation units (#5294). Nothing changes if the macro is not defined. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Mention JSON_NO_UDLS in the list of exported module symbols Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the user-defined string literals to <nlohmann/json_literals.hpp> Following the review in #5294, the literals now live in their own header instead of being removed entirely: <nlohmann/json.hpp> includes it at the end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined, so a project can opt out globally and include the header only where the literals are used. The header only uses public and standard macros, because the library's internal macros are undefined at the end of json.hpp and the amalgamation inlines macro_scope.hpp only once. For the same reason, the library no longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition is still visible to the header. The single-header copy is identical to the multi-header one, as it only includes <nlohmann/json.hpp>. The module always exports the literals. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test - Suppress clang-tidy misc-header-include-cycle on the intentional mutual include of json.hpp and json_literals.hpp. - Use operator"" _json with a space for GCC 4.8 in the test's detection aliases, as the header does. - Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Declare the literal operators through a local macro The GCC 4.8 spacing condition was repeated for both operator definitions and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix) now selects operator""##suffix or operator"" suffix in one place and is undefined at the end of json_literals.hpp. Suggested by gregmarr in review. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
66877675b1 |
Check the iterator range for binary values in basic_json(first, last) (#5719)
basic_json(first, last) treated value_t::binary like the structured types (array, object) in the range check, so it always copied the whole binary value regardless of the iterators, even for an empty range such as (b.end(), b.end()). The other primitive types (number, boolean, string) already reject such a range with invalid_iterator.204, and erase(first, last) already does the same for binary values, so this made the constructor inconsistent with both. Move case value_t::binary into the group of checked primitive types. Also update the two matching passages in basic_json.md that describe overload 7, and add a version-history note. Fixes #5670. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
7fd6895788 |
Hide a discarded container's content from the parser callback (#5706)
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.
The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.
Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.
Fixes #5643.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
6ae17630a4 |
Reject integral keys for contains(), find(), and count() at compile time (#5705)
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is a null pointer constant, so it converts to a null const char*, and the overloads taking const typename object_t::key_type& accepted it by constructing a std::string from that null pointer, which is undefined behavior (a crash with both libc++ and libstdc++). value(0, default_value) had the same problem in C++11, where the object comparator is not transparent. Add deleted overloads for integral arguments to contains(), find() (const and non-const), count(), and value() so that these calls are compile errors in every supported language mode instead of crashing. Calls with string, string_view, json_pointer, and size-typed element access (at(), operator[](), erase()) are unaffected. Fixes #5657. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
5bd766aa50 |
Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415) in write_bson_binary(), which only has the binary_t, not the basic_json value that holds it, so the exception was created with no JSON_DIAGNOSTICS context even though the equivalent to_msgpack() check names the value's path. The check also ran after the document size, all preceding elements, and this element's header and length had already reached the output adapter, so a caller-provided std::vector or std::string ended up holding a truncated document. calc_bson_sizes() already walks every value before anything is written, to size embedded documents and arrays and to reject invalid keys (out_of_range.409) up front. The subtype check now runs there instead, in calc_bson_binary_size(), which is given the basic_json value so the exception can use it as context. The now-redundant check in write_bson_binary() is removed, since calc_bson_sizes() always throws first if any binary value in the document has an oversized subtype. Fixes #5675. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
4bb1b14b06 |
Trim the compiler-appended NUL from wide/UTF string literals too (#5702)
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16, UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]")) failed with parse_error.101 at the terminating NUL of the literal, and accept() returned false. The array overload of input_adapter() only dropped the compiler-added trailing '\0' for arrays of char, so for wchar_t, char16_t, char32_t, and char8_t arrays that terminator was passed to the parser as data, which the macro then rejected. Broaden the trimming to every character type that a string literal can use (char, wchar_t, char16_t, char32_t, and, since C++20, char8_t). Arrays of any other element type (unsigned char, std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a trailing zero byte there is still read as genuine data. Fixes #5658. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bfe0f32d71 |
Fix std::terminate and null pointer access in input_stream_adapter (#5699)
Parsing from a std::istream crashed in two unusual but valid stream states, both in input_stream_adapter: - With eofbit in the stream's exceptions() mask, get_character() sets eofbit via is->clear(), which throws std::ios_base::failure. While that exception unwinds, ~input_stream_adapter() called clear() again to reset eofbit, which is still set and still in the exception mask, so it throws a second time out of the (implicitly noexcept) destructor and std::terminate() is called. The destructor now only calls clear() if a bit other than eofbit remains set, so the first exception can propagate normally. - For an std::istream without a stream buffer (rdbuf() == nullptr, e.g. std::istream(nullptr)), the constructor stored the null pointer without checking it, and get_character() dereferenced it. input_adapter(std::istream&) now throws parse_error.101 for such a stream, the same as it already does for a null FILE* or char*. Added regression tests to unit-deserialization.cpp and, for the JSON_PRECISE_STREAM_POSITION variant of get_character(), to unit-precise-stream-position.cpp; both crashed before this fix. Documented the two exceptions in parse.md and operator_gtgt.md. Fixes #5646. Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5b11a0282c |
Exchange the CustomBaseClass subobject in basic_json::swap() (#5697)
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap overload that forward to it) only exchanged m_data.m_type/m_data.m_value, leaving each value's json_base_class_t subobject in place. This is inconsistent with the copy and move constructors and copy assignment, which all carry the base class along with the value, so after a.swap(b) any metadata stored in a CustomBaseClass ended up attached to the wrong value. Algorithms that mix swap() with moves, such as std::sort, scrambled the metadata across the whole container. Fix the member swap() to also exchange the json_base_class_t subobject and extend the noexcept specifications of swap() and the friend swap() accordingly. Fixes #5653. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bfea6f36d3 |
Fix to_msgpack() reading the inactive number union member (#5694)
* Fix to_msgpack() reading the inactive number union member basic_json stores number_integer and number_unsigned in a union, and number_unsigned_t only has to be at least as wide as number_integer_t (with the default types, both are 64-bit and have the same representation). When number_integer_t is narrower, write_msgpack() read the wrong union member in two places: - The number_unsigned case wrote number_integer's bits instead of number_unsigned's, silently writing the wrong value whenever it did not fit in number_integer_t. - The number_integer case (non-negative branch) picked the encoded width by comparing number_unsigned's bits, which is undefined behavior, though the value written was still number_integer's, so at worst a too-wide encoding was chosen. Read the active member in both cases, like the other binary writers (CBOR, UBJSON, BJData, BSON, BON8) already do. Fixes #5644. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Cast number_integer to number_unsigned_t only once in to_msgpack() Addresses review comment by @gregmarr. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e444a66276 |
Copy values before inserting an initializer list into an array (#5693)
* Copy values before inserting an initializer list into an array
insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).
Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.
Fixes #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the reserve_array helper in the initializer_list insert fix
The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.
Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
6218aa212b |
Throw std::length_error for operator[](SIZE_MAX) instead of corrupting the array (#5687)
For idx == SIZE_MAX, the non-const array operator[] computed the new size as idx + 1, which wraps to 0. resize(0) then emptied the array, and the subsequent operator[](idx) on the now-empty vector wrote one element before its buffer. Every other too-large index (e.g. SIZE_MAX - 1) already went through resize(), which throws std::length_error and leaves the array unchanged; SIZE_MAX was the one value for which the overflow bypassed that safety net. Add a guard that throws std::length_error before computing idx + 1 when idx is the largest representable size_type value, so the array is left unchanged, matching the exception vector::resize() already throws for smaller (but still too large) indices. Fixes #5647. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
fc4c9c3446 |
Fix clear() to also reset the subtype of a binary value (#5680)
* Fix clear() to also reset the subtype of a binary value clear() on a binary value cleared the bytes but left the subtype untouched, so the result was not equal to a default-constructed binary value even though the documentation says clear() has the same effect as *this = basic_json(type()). The fix calls byte_container_with_subtype::clear_subtype() alongside the existing clear() call. Extended the "filled binary" clear() test in unit-modifiers.cpp with a case that uses a subtype, since the existing cases only covered binary values without one. Fixes #5669. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the table alignment in clear.md Addresses review comment by @gregmarr. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
65af260117 |
Move instead of deep-copy ordered_json values when an object grows (#5609)
* Move instead of deep-copy ordered_json values when an object grows ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>. With a std::string key, that pair is not nothrow move constructible (the const key has to be copied), so std::vector copies every element when it reallocates. For ordered_json, this deep-copies every member value an object already holds, including whole nested subtrees, on each growth step. Grow the storage in ordered_map instead, copying the keys and moving the values. This happens in two phases, so the strong exception guarantee is kept without try/catch. The first phase may throw, but only touches a temporary buffer: it copies the keys, value-initializes the values, and constructs the new element. The second phase moves the values (noexcept) and swaps the buffers. Because the new element is constructed before any value is moved, arguments that refer to elements of the container stay valid, as with std::vector. Types that cannot take this path keep the std::vector behavior. Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3): twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 -> 177 ms, canada unchanged. The number of allocations for twitter and citm_catalog drops by two thirds. Also add ParseStringOrdered rows to the benchmarks, and document the growth behavior and the exception safety of ordered_map. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: skip std::pair noexcept assumptions on EDG-based compilers ci_icpc and ci_nvhpc failed to compile unit-ordered_map.cpp: the static assertion that std::pair<const std::string, ordered_json> is not nothrow move-constructible fails there. The EDG front end (Intel icpc 2021.10, NVIDIA nvc++ 25.5) considers the defaulted move constructor of std::pair<const Key, T> noexcept even if copying Key can throw. With these compilers, std::vector already moves such elements itself when it grows, and ordered_map correctly leaves growing to it. The same misjudgement makes std::vector call std::terminate when a key copy throws during growth, so the exception-safety test with throwing_key would abort on these compilers as well. Skip the static assertion and the exception-safety section when __EDG__ is defined. Verified with icpc 2021.10 (-std=gnu++11) and nvc++ 25.5 (C++11 and C++17) on Compiler Explorer: unit-ordered_map and unit-disabled_exceptions build and pass. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d268eaa693 |
Convert floats with Eisel-Lemire when std::from_chars is unavailable (#5617)
* Convert floats with Eisel-Lemire when std::from_chars is unavailable Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit coordinates of canada.json) went to strtod unless std::from_chars was available. It is not used in C++11/14, and not with libc++, which does not define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's compute_float) now converts them with integer arithmetic, correctly rounded for any token with at most 19 significant digits. Longer tokens are truncated; the result is used if w and w + 1 round alike, else strtod decides as before. Overflow still yields infinity (out_of_range.406). The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit test recomputes every entry with big-integer arithmetic. Further tests: known values generated with Python (whose float() is correctly rounded), 200,000 round trips through to_chars, and the 128-bit multiplication and leading-zero count against big-integer references (both with and without a 128-bit type). Checked against strtod on 6.5 million tokens, among them 60,000 exact halfway cases: no difference. json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang); other files unchanged. Compile time of a TU including json.hpp: +0.7%. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: unused parse result and find() == npos in the Eisel-Lemire tests GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...)) because parse() is [[nodiscard]]; assign the result to a dummy json as the other tests do. clang-tidy flagged longer.find('.') == npos with abseil-string-find-str-contains; store the position in a variable first. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9e1a09eec0 |
Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings, so such maps are rejected. The error so far was the one for a malformed string (e.g. "expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not tell the user what went wrong. Report the type of the key instead: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0 The exception id (parse_error.113) and type are unchanged. Malformed string keys and a missing key keep their previous messages. Document the restriction on the CBOR and MessagePack pages. Refs #2766, #3381 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Point the MessagePack key note to the spec's profile section The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
56ddcb65f0 |
Fix documentation style check: wrap long line in cbor_tag_handler_t.md (#5603)
The line added in #5559 exceeded the 160-character limit enforced by docs/mkdocs/scripts/check_structure.py, breaking the documentation build. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1e101ecac1 |
Add BON8 support (#2998)
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6bd106893a |
Fix CBOR tag handling in cbor_tag_handler_t::store for non-binary items (#5559)
When using cbor_tag_handler_t::store, tags 0xD8-0xDB previously assumed that the tagged item was a byte string, unconditionally attempting to parse binary data and failing on valid CBOR documents containing tags applied to integers, strings, arrays, or objects (such as self-describe tag 55799). Check whether the tagged data item is a byte string (0x40-0x5B or 0x5F). If it is a byte string, store the subtype on the binary value as before. Otherwise, iteratively process the tagged value in the driver loop using item_read so that chained tags do not consume native stack space. Part of #5316. Signed-off-by: ReturnKartikey <kartikeynegi2000.work@gmail.com> |
||
|
|
f7972970a4 |
Throw instead of writing MessagePack lengths beyond UINT32_MAX (#5584)
* Throw instead of writing MessagePack lengths beyond UINT32_MAX MessagePack stores the length of a string, binary value, array, or object in at most 32 bits. For a larger value, to_msgpack wrote no length at all, so the output could not be read back. It now throws out_of_range.412, which BSON already uses for its 32-bit length fields. The check lives in one function, so each length is written by an if/else chain that ends in a plain else, without a condition that can never be false. It is tested with string and binary types that report a size beyond UINT32_MAX without allocating it, like the BSON tests do. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the CI failures of the MessagePack length check - mark to_msgpack_length's value as used when exceptions are disabled (-Wunused-parameter, misc-unused-parameters) - put "Exception safety" before "Exceptions" in to_msgpack.md, as the documentation style check requires - create the test's string value from its type: constructing it from a beyond_uint32_string_t considers the std::filesystem::path conversion, which libstdc++ 10 reports as ambiguous for a class derived from std::string (clang 13) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip the MessagePack string length test for clang with libstdc++ 10 C++17 builds consider the std::filesystem::path conversion for the string type, and with clang and libstdc++ 10 that conversion is ambiguous for a class derived from std::string. Creating the value from its type did not avoid it, since any basic_json with that string type instantiates the check. The binary and ext cases are still tested there. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Keep the MessagePack string test type and its alias in one block astyle indented the alias oddly when it had an #ifdef of its own after the binary alias; declare it right after the string type, in the same block. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
95e9a5931c |
Write BSON in linear time, without recursing per nesting level (#5553)
* Write BSON in linear time, without recursing per nesting level to_bson() had two problems with nested values: - It recursed once per nesting level, so a value nested deeply enough - 100,000 levels on an 8 MiB stack - exhausted the call stack and terminated the process, although parse() accepts such values without complaint. - BSON prefixes every document and array with its length. The writer computed that length by walking the entire value below it, again for every nested document it wrote, which made serializing O(size x depth). A 200-level document took 30 ms instead of 1. Both passes are now iterative, and each length is computed exactly once: - calc_bson_sizes() computes the length of every document and array in one pass, each from the lengths of its entries, into a table ordered the way they are written. - write_bson_document() then writes the document, taking each length from the table. Everything observable is unchanged, as a differential test against develop confirms byte for byte: - The same bytes are written. - A key containing U+0000 still throws out_of_range.409 for the same first key, with the same diagnostics path, before anything is written. - A document too large for BSON still throws out_of_range.412 before anything is written. - A binary subtype above 255 still throws out_of_range.415 after the same partial output. Only the enclosing objects and arrays are kept on a stack, so a flat document allocates nothing for it. Measured against develop (clang -O3, median of 201 runs): flat objects unchanged, flat arrays 37% faster (the array length was computed twice), a nested 3,000-object document 2x faster, a 200-level document 33x faster. to_bson.md documented the quadratic complexity since #5334; it is linear again. Fixes #5392 for BSON, and #5308. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Do not require a default-constructible string_t in the BSON writer GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default constructor; develop never default-constructed string_t here either. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Let the BSON index-name helper only fill its output parameter It returned a reference to the string it filled, so callers held a second name for index_name. Addresses review feedback. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1e44262091 |
Make JSON_STRICT_NUL_HANDLING part of the ABI tag (#5560)
* Make JSON_STRICT_NUL_HANDLING part of the ABI tag JSON_STRICT_NUL_HANDLING (#5534) changes the bodies of inline functions: the lexer's handling of '\0' and input_adapter() for char arrays. So translation units compiled with and without it define the same functions differently, an ODR violation - the case the ABI tag exists for, as with JSON_BRACE_INIT_COPY_SEMANTICS (_bics). It now appends _snul to the inline namespace. The macro is new in 3.13.0, so no existing namespace changes. Its default moves to abi_macros.hpp, and it is only #undef'd without JSON_TEST_KEEP_MACROS, as for the other ABI macros. The ABI config tests, the namespace docs, the macro's docs and the Natvis file cover the new tag. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
c60a0bc336 |
Allocate the deep copy's key scratch space with the provided allocator (#5573)
* Allocate the deep copy's key scratch space with the provided allocator The iterative deep copy builds each object's keys in a temporary vector of key/value pairs before handing them to the object's range constructor. That vector holds basic_json values, so like the values themselves it now uses AllocatorType instead of std::allocator. Also document that AllocatorType covers the JSON values, while most temporary storage still uses std::allocator. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Count allocate_at_least in the scratch-counting test allocator From C++23 on, libc++'s containers allocate through allocate_at_least when the allocator has one. The test allocator inherited it from std::allocator, so the scratch allocations were not counted and the test failed on Xcode. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d19f7f5dce |
Fix BSON conformance issue (#5185)
* 🐛 fix BSON conformance issue Signed-off-by: Niels Lohmann <mail@nlohmann.me> * 🐛 fix BSON conformance issue Signed-off-by: Niels Lohmann <mail@nlohmann.me> * 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531) from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded text string into the resulting json value without any UTF-8 validation, even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications all require text strings to be valid UTF-8. Malformed input only failed later, if the value was dump()'d, with a type_error.316 - so the allow_exceptions=false pattern used specifically to get a discarded sentinel instead of an exception did not discard this category of malformed input, unlike every other kind of malformed binary input this library rejects at decode time (see #5529). Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON string reads, binary_reader::get_string(): validate the bytes with the UTF-8 DFA right after they are read, and report failures the same way as every other binary_reader error (parse_error.113), so allow_exceptions and strict discarding behave consistently. get_binary()/binary blob reads are untouched and still accept arbitrary bytes, since only text strings are required to be UTF-8. There were two independent implementations of a UTF-8 validator: the lexer's streaming scanner, and the serializer's Hoehrmann DFA used by dump_escaped_impl(). Rather than write a third, the serializer's decode() function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are extracted into detail/string_utils.hpp (a low-level header already included before both detail/input/ and detail/output/), alongside a new is_valid_utf8() helper built on the same decode() step. serializer.hpp's dump_escaped_impl() now calls the shared decode(), so there is exactly one UTF-8 validator in the codebase; dump()'s exact type_error.316 messages and byte-index reporting are unchanged (see the added regression-guard test in unit-serialization.cpp). Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * ⚡ validate only newly read bytes of binary-format strings get_string() validated the whole result after each call, but get_bytes() appends to it and CBOR indefinite-length strings collect all chunks in the same result, so every chunk re-validated everything read before it. An input of many small chunks took quadratic time (80000 one-byte chunks, 160 KB of input, took about 7 seconds). Only the newly read bytes are validated now, which also matches RFC 8949's requirement that every chunk is valid UTF-8 on its own. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
632a5812a8 |
Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041) (#5272)
* Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041) NLOHMANN_DEFINE_TYPE_INTRUSIVE(Type) and its 11 sibling macros produced broken code for types with no members to serialize. Invoking a variadic macro so __VA_ARGS__ is empty is only standard-conforming since C++20, so a plain __VA_OPT__ fix (as tried in #5142) breaks every pre-C++20 build under -pedantic. Instead, make all 12 macros purely variadic and dispatch on argument count using a sentinel-padded extension of the existing NLOHMANN_JSON_GET_MACRO idiom, giving full C++11-C++26 support with no feature-test gate. Verified against real GCC 16 and Clang at -std=c++11/14/17/20 with -pedantic -Werror -Wvariadic-macros: zero regressions in the existing unit-udt_macro.cpp suite plus 12 new zero-member test cases. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI failures in zero-member NLOHMANN_DEFINE_TYPE_* macros Three issues surfaced on PR #5272's real CI that weren't caught by local testing against a narrower flag set: - GCC -Werror=noexcept: the four truly-empty from_json bodies (plain INTRUSIVE/NON_INTRUSIVE, with and without _WITH_DEFAULT) provably never throw but weren't declared noexcept; mark them noexcept explicitly. to_json and the derived-type from_json overloads are left alone since they genuinely can throw (object assignment / delegating to the base class's from_json). - clang-tidy bugprone-macro-parentheses: false positive on the same 8 zero-member bodies (Type/BaseType used purely as declarator types); suppressed with NOLINTNEXTLINE comments in the same style already used elsewhere in this file (see NLOHMANN_JSON_SERIALIZE_ENUM). - MSVC's traditional preprocessor doesn't fully expand NLOHMANN_JSON_CAT(prefix, NLOHMANN_JSON_TYPE_TAG(...))(...) in one pass, which broke a pre-existing one-member usage in unit-regression2.cpp with syntax errors. Wrap all 12 public dispatcher macros in an extra outer NLOHMANN_JSON_EXPAND(...), matching the pattern NLOHMANN_JSON_PASTE already uses for the same MSVC quirk. Re-verified against real GCC 16 and Clang at -std=c++11/14/17/20 with -pedantic -Werror -Wvariadic-macros -Wnoexcept, including the exact files that failed in CI (unit-udt_macro.cpp, unit-regression2.cpp), against both the modular headers and the re-amalgamated single header. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix clang-tidy misc-const-correctness in unit-udt_macro.cpp The four zero-member ONLY_SERIALIZE test objects are only ever read (via to_json), never mutated, so mark them const per clang-tidy. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix derived-type macro dispatch capping members at 62 instead of 63 NLOHMANN_JSON_GET_MACRO resolves 64 positional arguments, with NAME at position 65. NLOHMANN_JSON_TYPE_TAG dispatches on Type plus the member list, so it resolves correctly up to the 63 members NLOHMANN_JSON_PASTE supports. NLOHMANN_JSON_DERIVED_TYPE_TAG dispatched on the two-token Type,BaseType prefix plus the member list, running out one slot early: at 63 members, position 65 landed on the last member name instead of a sentinel and NLOHMANN_JSON_CAT built an undefined identifier such as NLOHMANN_JSON_DEFINE_DERIVED_TYPE_INTRUSIVE_m63, with the compiler reporting "unknown type name 'm1'" once per member and nothing pointing at an argument-count limit. That silently reduced all six NLOHMANN_DEFINE_DERIVED_TYPE_* macros from 63 members to 62, contradicting the "up to 63 members" contract in docs/mkdocs/docs/api/macros/nlohmann_define_derived_type.md. Drop the leading Type and defer to NLOHMANN_JSON_TYPE_TAG so the tag is computed from BaseType plus the member list, which fits the available slots. The zero-own-member derived bodies are therefore selected by tag 1 rather than 2, and the sentinel table for the derived tag is no longer needed. Add a regression test at the documented maximum for both the plain and the derived macros; it fails to compile against the previous dispatch. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the zero-member macro bodies by intent, not argument count The dispatch tag was the literal token 1 or N, pasted onto a macro prefix to select the zero-member or member-carrying body. For the derived-type macros that reads wrong: their tag is computed after dropping the leading Type, so the zero-member body was named _1 while taking two parameters (Type, BaseType). Emit EMPTY and MEMBERS instead. The mechanism is unchanged -- the tag is still a token pasted onto the prefix by NLOHMANN_JSON_CAT -- but the body names now say what they are rather than encoding an argument count that only lines up for half of the macros. Collapse the four duplicated zero-member bodies while here: with no members there is nothing to default, so each _WITH_DEFAULT_EMPTY body was a byte-for-byte copy of its plain counterpart. They are now one-line aliases, leaving a single definition of what an empty object serializes to per intrusive/non-intrusive and base/derived combination. No functional change: for both zero-member and member-carrying types the preprocessed to_json/from_json output is token-for-token identical, and the arity limits are unchanged (63 members, base and derived). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document zero-member support in the macro API reference docs/mkdocs/docs/features/arbitrary_types.md already gained a note, but the three api/macros pages are where the parameter contract is actually specified and they still described member as a non-empty list. State that the list may be empty on each page, and add a note showing what the zero-member case generates: an empty JSON object for the plain macros, and base-type-only serialization for the derived ones. Both notes record that the WITH_NAMES variants do not support this. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Keep user macros named EMPTY or MEMBERS out of the member-count dispatch The dispatch produced the bare token EMPTY or MEMBERS and pasted it onto the macro prefix afterwards. In between, the token was rescanned, so a user macro with either name replaced it: with `#define MEMBERS x` in scope, even NLOHMANN_DEFINE_TYPE_INTRUSIVE(A, member) -- which compiled before -- expanded to garbage, and `#define EMPTY` broke the zero-member form. Paste the suffix onto the prefix directly in the GET_MACRO slot table instead. Operands of ## are not macro-expanded, so the selected body name is formed before any user macro can interfere. NLOHMANN_JSON_TYPE_TAG and NLOHMANN_JSON_DERIVED_TYPE_TAG become NLOHMANN_JSON_TYPE_BODY and NLOHMANN_JSON_DERIVED_TYPE_BODY, taking the prefix as their first argument; NLOHMANN_JSON_CAT is no longer needed. The body macro names are unchanged, and so is the generated code. Add a regression test that defines EMPTY and MEMBERS around plain and derived types, with and without members. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test for EMPTY and MEMBERS so -Wunused-macros accepts them Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
ad2c14b985 |
Document response times, supported versions, access, secrets, and dependency policies (#5580)
Answer the OpenSSF Best Practices criteria that asked for policies the project follows but had not written down: - SECURITY.md: a first response within 14 days, publishing an advisory with credit once a fix is released, and that only the latest release receives security fixes. - Governance: who has access to the project's resources, how write or admin access is granted, and how CI secrets are stored and rotated. - Quality assurance: how dependencies of the build, test, and documentation tooling are pinned, scanned, and kept free of known vulnerabilities. Also update the assurance case, since comparison no longer recurses per nesting level (#5390), and point the best practices badge and links to bestpractices.dev under the program's current name. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
465407f3ce |
Improve error message for const fields (#2818)
* Improve error message for const fields * Reject const arguments to get_to() with a clear message Reword the static_assert, add it to the C array overload of get_to() as well, and document that v must not be const. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
02dd3e67f2 |
Fix stack overflow and exponential runtime when comparing nested values (#5390)
* Compare values without recursing, and without comparing them twice Comparing two values compared their containers, which compare their elements, which brought the comparison back once per nesting level. Two values nested deeply enough exhausted the call stack and terminated the process with a segmentation fault - the same bug as #5387, in the last operation that still had it. Worse, an ordered comparison took exponentially long in the nesting depth before C++20. std::vector's operator< is a lexicographical comparison, which asks whether an element is less than its counterpart and then whether the counterpart is less than it - two full comparisons of everything below that element, at every level. Comparing two equal values nested 30 levels deep, which is nothing unusual, took 3.8 seconds; 40 levels would have taken an hour, and nothing about the value has to be pathological to get there. C++20 is unaffected: std::lexicographical_compare_three_way asks once. Compare a value that is nested too deeply to descend into on an explicit stack instead, in a single pass that yields less, equal, greater or unordered at once. Equality and the three-way comparison descend as they always did for the first 128 levels, which nothing measurable costs them; an ordered comparison no longer descends at all, which is what takes the exponent out of it. Objects and arrays that are not nested deeply are otherwise compared exactly as before. The results are unchanged for every pair of values: 68121 comparisons of a corpus that covers NaN, discarded values, mixed number types, binary values, empty containers and both object types are identical to develop, in C++11, C++17 and C++20, with and without thread_local storage and legacy discarded comparison. Reproducing that meant reproducing two subtleties: a lexicographic comparison steps over a pair it cannot order, where a three-way comparison stops at it, and an object compares its keys with < where its entries are ordered but with == where they are only checked for equality - not with the object's own comparator, which for nlohmann::ordered_map tells equality. Equality needs no ordering, so it no longer asks for any: a key or string type that can only be compared for equality still works. Measured (medians of 7 interleaved runs, clang -O3, C++11): comparing two equal values nested 30 levels deep 3778 ms -> 0.002 ms; ordering flat objects -33.6%; ordering flat arrays of numbers +27.3%, the one shape that pays for the single pass; equality unchanged throughout. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Describe comparison in the no-thread-local docs and CI target Comparing two values now bounds its descent with a thread_local counter just as copying does, so the JSON_NO_THREAD_LOCAL page, the macro overview and the ci_test_no_thread_local target cover both rather than copying alone. Also record what switching the macro on costs a comparison: on the benchmark documents, comparing two equal values takes 10% to 90% longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Take the descent flag as an argument rather than testing it MSVC reports the test of a constant as C4127 ("conditional expression is constant"), which the Windows builds treat as an error: may_descend is false for operator<, so the operand short-circuits the whole condition. Passing it to compare_descent_exhausted() puts the test where the value is an ordinary parameter, and leaves the call sites with no condition of their own. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Note the comparison fallback in the no-thread-local documentation The macro page describes what the library defines JSON_NO_THREAD_LOCAL for by itself in terms of copying alone; comparing falls back the same way. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Parenthesise the reserve() computation in the comparison test clang-tidy reports the mixed * and + as readability-math-missing- parentheses, as it does for the identical line in the copy test. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use the shared descent bookkeeping rather than a second set Comparing kept a thread_local count, a limit and a guard of its own beside the ones copying already had, all three the same thing under a different name. They are gone; the shared count, limit and guard do the work. The guard grows a second constructor here, because the comparison operators are written as a macro and a macro cannot use the preprocessor: it cannot look the count up behind an #ifdef the way copy_structured does, so the guard looks it up for it. nesting_depth_exhausted() arrives for the same reason - whether an operator descends at all is a constant at every call site, and testing it there is what MSVC reports as C4127. Also say in compare_leaves what happens to a pair that is an array on one side and an object on the other, since the answer is not obvious from the code: an operator only descends into two values of the same type, so such a pair is told apart by its types alone - unequal, and ordered the way the types are - exactly as it is above the bound. And record what the explicit stack costs: the comparison operators are noexcept and the container comparison this replaces allocated nothing, so running out of memory here ends the process instead of throwing. It takes a value nested past the bound and an exhausted heap to reach, and the same comparison used to exhaust the call stack, but it is a new way to fail. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
abbe52d6de |
Add JSON_PRECISE_STREAM_POSITION to leave the character that terminates a number in the stream (#5344)
* docs: qualify the operator>> stream positioning guarantee operator>>'s notes state that it leaves the stream positioned right after the parsed value, so that concatenated JSON values can be read back to back. That does not hold when the value is a number: a number is only terminated by the character that follows it, and the lexer's unget() is simulated (it rewinds only the lexer's own bookkeeping), so that character stays consumed from the stream. Document the actual behaviour: the guarantee holds for all value types except numbers, which must be followed by whitespace. Also qualify the cross-reference on the JSON Lines page, which repeated the unqualified claim. Documentation only; the behaviour itself is tracked in #5340. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * fix: restore the character that terminates a number (#5340) operator>> is documented to leave the stream positioned right after the parsed value, so that concatenated JSON values can be read back to back. That did not hold for numbers: a number is only terminated by the character following it, and lexer::scan_number() reads that character and calls unget() -- which is simulated and rewinds only the lexer's own bookkeeping. input_stream_adapter consumes via sbumpc() with no matching sungetc(), so the terminating character stayed consumed and the next extraction started one byte too late ('1true' left the stream at 'rue'). Propagating unget() to the adapter directly does not work: next_unget makes the following get() replay the cached character, so the terminator would be delivered twice. Instead, restore the still-pending character once at the end of a non-strict parse, where the input is handed back to the caller: - input_stream_adapter gains unget_character() (sungetc()) and advertises it via supports_unget, detected the same way as supports_seek. - lexer::restore_pending_unget() turns a pending simulated unget of a real (non-EOF) character into a real one and clears next_unget so the character is not also replayed. It is a no-op for adapters that cannot unget, and reports failure when sungetc() fails, in which case the input is left as it was before. - parser calls it on the three non-strict paths, i.e. for operator>> and sax_parse(strict = false). Strict parse()/accept() are unaffected: they require the input to end after the value, so the character is consumed by the end-of-input check anyway. Parse error messages and reported positions are unchanged. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * tests: fix CI failures in the #5340 test helpers Four CI failures, all in the new test code: - GCC (-Werror=useless-cast): drop the `json(...)` wrapper around `json::parse(...)`, which already returns a `json`. - GCC (-Werror=unused-result): assign the discarded `json::parse()` result to a dummy, the idiom used elsewhere in the test suite, and catch `json::parse_error&` for consistency. - clang-tidy (google-default-arguments): remove the default argument from the `pbackfail()` override; `sungetc()` supplies the base declaration's default. - MSVC (bad allocation): `no_putback_streambuf::underflow()` set a one-character get area without advancing `m_pos`, so an implementation whose `istream::get` peeks before it bumps re-read the same character forever. Keep no get area at all: `underflow()` peeks, `uflow()` consumes, and `sungetc()` still always lands in `pbackfail()`, which is what the test needs. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * fix: leave the character that terminates a number in the input Read the character following a number without consuming it, instead of consuming it and putting it back. input_stream_adapter now peeks with sgetc() and only steps over the character when the next one is requested or when the adapter is destroyed, so releasing it cannot fail - no putback position is required from the streambuf. Suggested by gregmarr in #5344. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * docs: match the version history wording to the peek-based fix Signed-off-by: Niels Lohmann <mail@nlohmann.me> * docs: drop the whitespace-separator caveat from the parsing pages The caveat added in #5343 describes the behavior this branch fixes: a number no longer consumes the character that terminates it, so concatenated values need no separator. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * refactor: split the strict and non-strict paths in parser Folding the release_lookahead() call into the existing strict check left the "in strict mode" comment on an else-if branch, and made the strict condition in sax_parse() redundant with the branch it followed. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Put the stream position fix behind JSON_PRECISE_STREAM_POSITION Leaving the character that terminates a number in the stream is observable: reading "1,2,3" with repeated operator>> works today only because the comma after each number is swallowed, and std::getline after a number skips the line break. Both break with the fix, so make it opt-in for 3.x, as suggested by @gregmarr in the review. - JSON_PRECISE_STREAM_POSITION (default 0) selects the peek-based input_stream_adapter. Without it, the adapter is the consuming one from develop and has no supports_lookahead, so lexer::release_lookahead() and the parser's calls to it compile to nothing. - The macro changes input_stream_adapter's layout and member functions, so it gets the ABI tag _psp, after _bics. The ABI config tests, the natvis generator, and nlohmann_json.natvis (regenerated) know the tag. - The tests for the fix move to unit-precise-stream-position.cpp, which defines the macro itself and runs in every build, and gain the two cases above. unit-deserialization.cpp pins the default behavior instead. - The docs describe the default behavior again and point to the new macro page; version history says "added in 3.13.0, planned default in 4.0.0". Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |