mirror of
https://github.com/nlohmann/json.git
synced 2026-10-01 22:45:17 +00:00
* Share scalar serialization between dump_internal and dump_value dump_value()'s cases for string, binary, boolean, number_integer, number_unsigned, number_float, discarded and null were a byte-for-byte copy of dump_internal()'s (added together in #5285 for the iterative fallback path). Any future change to scalar output had to be made in both places, or the recursive and depth-limited paths would silently start producing different bytes. Extract the shared cases into a private dump_scalar() and have both dump_internal() and dump_value() call it. Output is unchanged: dump(), dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t variants are byte-identical over the json_test_data corpus before and after, and dump() throughput on a scalar-heavy document is unaffected. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Drop serializer.hpp's dependency on binary_writer.hpp The only use of binary_writer in serializer.hpp was binary_writer<BasicJsonType, char>::to_char_type() to write the U+FFFD replacement character's three bytes. With CharType=char this is an identity conversion, so the include of binary_writer.hpp (and transitively binary_reader.hpp) pulled in a large, unrelated header for a no-op call. Write the three bytes directly instead. serializer.hpp compiles standalone with -Wall -Wextra -Werror, with and without -funsigned-char, and unit-serialization's error_handler_t::replace cases (with and without ensure_ascii) still pass. Moving binary_writer's to_char_type/to_msgpack_length to its private section is left as an optional follow-up. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale #include lines in the output headers output_adapters.hpp included <algorithm> and <iterator> for std::copy and std::back_inserter, which have not been used there since #3569 (2022). serializer.hpp included <algorithm> for std::reverse (also unused), <cmath> for labs/isnan/signbit (only std::isfinite is used) and <utility> for std::move (nothing from <utility> is used there), while using std::next without including <iterator> at all, relying on getting it transitively through output_adapters.hpp's own stale <iterator>. Drop the unused includes, add <iterator> for std::next, and correct the remaining include comments. Both headers still compile standalone with -Wall -Wextra -Werror. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides output_vector_adapter, output_stream_adapter and output_string_adapter declared their write_characters(const CharType*, std::size_t) override JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with a null pointer and length 0 for an empty string or binary value; the type-erased call path only stayed silent under UBSan because the static callee at those call sites is the unattributed virtual base. A nonnull attribute on a definition lets GCC and Clang assume the parameter is non-null inside the function body even when the call is virtual, so this was latent undefined behavior, not just style. Drop the attribute from the three overrides and document the (nullptr, 0) contract on output_adapter_protocol::write_characters. unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise empty binary/string payloads through the stream and vector/string adapters) pass under -fsanitize=address,undefined,nonnull-attribute. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update stale serializer doc comments to match the current implementation dump_internal()'s doc block still described the pre-#5285/#5449 implementation: an escape_string() function that does not exist (the function is dump_escaped), integer conversion "implicitly via operator<<" (dump_integer actually uses a digit-pair lookup table), and floating-point conversion via "%g" (IEEE-754 types go through to_chars, others through snprintf). dump_value()'s comment said elements are pushed for dump_internal to walk, but it is dump_iteratively() that walks the stack. dump_escaped(), dump_integer() and dump_float() each said they write "to output stream @a o", which has not been true since the writer moved to write_buffer. Doc-only change; no behavior, API or ABI impact. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Merge duplicate byte-to-hex helper and drop stale '| 0' promotions serializer::hex_bytes() and binary_writer::hex_byte() had identical bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it from both. Also drop the `| 0` at the two serializer call sites (hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8) added it so that `ss << std::hex << (byte | 0)` printed a number rather than a char with the old stringstream writer; the int result just narrows back to uint8_t now, so it was a no-op. Behavior is unchanged: unit-serialization, unit-bon8 (whose type_error.316 messages exercise this code) and unit-diagnostics pass, and both headers still compile standalone with -Werror. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Trim two stale lint suppressions in serializer.hpp dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg, left over from the snprintf-based implementation (#3088); there is no variadic call on that line, so keep only the qualified-auto suppressions it actually needs. remove_sign()'s assert checked `x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a NOLINT(misc-redundant-expression) to hide it; the second conjunct is always true once x < 0, and has been since6ce2f35ba(2019), so reduce the assert to `x < 0` and drop the suppression instead of masking it. Both are documentation-only changes to assertions/suppressions, not behavior. The to_chars.hpp `#if 0` branch this item also flagged is left alone, next to draft PR #5634's pending hunk. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the Hoehrmann SPDX copyright line to string_utils.hpp serializer.hpp carried the SPDX-FileCopyrightText line for Björn Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d table) has lived in string_utils.hpp since #5185 (d19f7f5dc); serializer.hpp now only calls decode(). Move the copyright line to where the code it covers actually is. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop calling std::localeconv() on every dump() The serializer constructor snapshotted std::localeconv() into a locale_chars member on every dump(), even though the only reader is dump_float(number_float_t, std::false_type)'s snprintf path, taken only for a number_float_t that is neither IEEE single nor double. localeconv() is not required to be thread-safe with setlocale(), so every dump() paid for and raced on a lookup that almost never mattered. Remove locale_chars and the locale member. Right before the thousands-separator/decimal-point fixups in the snprintf path, read std::localeconv() into local thousands_sep/decimal_point variables (null-checked, first byte only, as before) - the same way lexer::get_decimal_point() already does since #5597. Output is unchanged unless the locale changes during a single dump(); in that case the fixups now match what snprintf just produced, instead of a value snapshotted before the call. Overlaps draft PR #5608, which touches the same constructor and dump_float() lines to move this code into a new dump_float_snprintf(); this lands the lookup change now as #5709 asks, and #5608 can do the lookup inside dump_float_snprintf() when it rebases. Verification: the full json_test_data corpus (742 files, dump(), dump(4) and dump(-1,' ',true)) is byte-identical to before the change under the C locale. Added a test pinning the new per-conversion lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that switches on its first write, after the serializer's write buffer has been flushed once but before a later float is converted) and checks the decimal point is still normalized using the locale active at conversion time. On a platform where long double is IEEE-754 double (e.g. 64-bit Arm), dump_float() takes the locale-independent to_chars() path and the test is a no-op there; it is meaningful on a platform where long double is extended precision (most x86 targets). Part of #5709 item 3 Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>