Commit Graph

7 Commits

Author SHA1 Message Date
Niels Lohmann
3926fcaac3 Deduplicate serializer dump code; fix stale includes, docs, and lint (#5729)
* Share scalar serialization between dump_internal and dump_value

dump_value()'s cases for string, binary, boolean, number_integer,
number_unsigned, number_float, discarded and null were a byte-for-byte
copy of dump_internal()'s (added together in #5285 for the iterative
fallback path). Any future change to scalar output had to be made in
both places, or the recursive and depth-limited paths would silently
start producing different bytes.

Extract the shared cases into a private dump_scalar() and have both
dump_internal() and dump_value() call it. Output is unchanged: dump(),
dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t
variants are byte-identical over the json_test_data corpus before and
after, and dump() throughput on a scalar-heavy document is unaffected.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop serializer.hpp's dependency on binary_writer.hpp

The only use of binary_writer in serializer.hpp was
binary_writer<BasicJsonType, char>::to_char_type() to write the
U+FFFD replacement character's three bytes. With CharType=char this
is an identity conversion, so the include of binary_writer.hpp (and
transitively binary_reader.hpp) pulled in a large, unrelated header
for a no-op call.

Write the three bytes directly instead. serializer.hpp compiles
standalone with -Wall -Wextra -Werror, with and without
-funsigned-char, and unit-serialization's error_handler_t::replace
cases (with and without ensure_ascii) still pass. Moving
binary_writer's to_char_type/to_msgpack_length to its private section
is left as an optional follow-up.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale #include lines in the output headers

output_adapters.hpp included <algorithm> and <iterator> for std::copy
and std::back_inserter, which have not been used there since #3569
(2022). serializer.hpp included <algorithm> for std::reverse (also
unused), <cmath> for labs/isnan/signbit (only std::isfinite is used)
and <utility> for std::move (nothing from <utility> is used there),
while using std::next without including <iterator> at all, relying on
getting it transitively through output_adapters.hpp's own stale
<iterator>.

Drop the unused includes, add <iterator> for std::next, and correct
the remaining include comments. Both headers still compile standalone
with -Wall -Wextra -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides

output_vector_adapter, output_stream_adapter and output_string_adapter
declared their write_characters(const CharType*, std::size_t) override
JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with
a null pointer and length 0 for an empty string or binary value; the
type-erased call path only stayed silent under UBSan because the
static callee at those call sites is the unattributed virtual base.
A nonnull attribute on a definition lets GCC and Clang assume the
parameter is non-null inside the function body even when the call is
virtual, so this was latent undefined behavior, not just style.

Drop the attribute from the three overrides and document the
(nullptr, 0) contract on output_adapter_protocol::write_characters.
unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise
empty binary/string payloads through the stream and vector/string
adapters) pass under -fsanitize=address,undefined,nonnull-attribute.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update stale serializer doc comments to match the current implementation

dump_internal()'s doc block still described the pre-#5285/#5449
implementation: an escape_string() function that does not exist
(the function is dump_escaped), integer conversion "implicitly via
operator<<" (dump_integer actually uses a digit-pair lookup table),
and floating-point conversion via "%g" (IEEE-754 types go through
to_chars, others through snprintf). dump_value()'s comment said
elements are pushed for dump_internal to walk, but it is
dump_iteratively() that walks the stack. dump_escaped(), dump_integer()
and dump_float() each said they write "to output stream @a o", which
has not been true since the writer moved to write_buffer.

Doc-only change; no behavior, API or ABI impact.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge duplicate byte-to-hex helper and drop stale '| 0' promotions

serializer::hex_bytes() and binary_writer::hex_byte() had identical
bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it
from both. Also drop the `| 0` at the two serializer call sites
(hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8)
added it so that `ss << std::hex << (byte | 0)` printed a number
rather than a char with the old stringstream writer; the int result
just narrows back to uint8_t now, so it was a no-op.

Behavior is unchanged: unit-serialization, unit-bon8 (whose
type_error.316 messages exercise this code) and unit-diagnostics
pass, and both headers still compile standalone with -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Trim two stale lint suppressions in serializer.hpp

dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried
NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg,
left over from the snprintf-based implementation (#3088); there is no
variadic call on that line, so keep only the qualified-auto
suppressions it actually needs. remove_sign()'s assert checked
`x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a
NOLINT(misc-redundant-expression) to hide it; the second conjunct is
always true once x < 0, and has been since 6ce2f35ba (2019), so
reduce the assert to `x < 0` and drop the suppression instead of
masking it.

Both are documentation-only changes to assertions/suppressions, not
behavior. The to_chars.hpp `#if 0` branch this item also flagged is
left alone, next to draft PR #5634's pending hunk.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the Hoehrmann SPDX copyright line to string_utils.hpp

serializer.hpp carried the SPDX-FileCopyrightText line for Björn
Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d
table) has lived in string_utils.hpp since #5185 (d19f7f5dc);
serializer.hpp now only calls decode(). Move the copyright line to
where the code it covers actually is.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop calling std::localeconv() on every dump()

The serializer constructor snapshotted std::localeconv() into a
locale_chars member on every dump(), even though the only reader is
dump_float(number_float_t, std::false_type)'s snprintf path, taken
only for a number_float_t that is neither IEEE single nor double.
localeconv() is not required to be thread-safe with setlocale(), so
every dump() paid for and raced on a lookup that almost never mattered.

Remove locale_chars and the locale member. Right before the
thousands-separator/decimal-point fixups in the snprintf path, read
std::localeconv() into local thousands_sep/decimal_point variables
(null-checked, first byte only, as before) - the same way
lexer::get_decimal_point() already does since #5597. Output is
unchanged unless the locale changes during a single dump(); in that
case the fixups now match what snprintf just produced, instead of a
value snapshotted before the call.

Overlaps draft PR #5608, which touches the same constructor and
dump_float() lines to move this code into a new
dump_float_snprintf(); this lands the lookup change now as #5709 asks,
and #5608 can do the lookup inside dump_float_snprintf() when it
rebases.

Verification: the full json_test_data corpus (742 files, dump(),
dump(4) and dump(-1,' ',true)) is byte-identical to before the change
under the C locale. Added a test pinning the new per-conversion
lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that
switches on its first write, after the serializer's write buffer has
been flushed once but before a later float is converted) and checks
the decimal point is still normalized using the locale active at
conversion time. On a platform where long double is IEEE-754 double
(e.g. 64-bit Arm), dump_float() takes the locale-independent
to_chars() path and the test is a no-op there; it is meaningful on a
platform where long double is extended precision (most x86 targets).

Part of #5709 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:31:42 +02:00
Niels Lohmann
d19f7f5dce Fix BSON conformance issue (#5185)
* 🐛 fix BSON conformance issue

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 🐛 fix BSON conformance issue

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)

from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).

Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.

There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).

Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* ⚡ validate only newly read bytes of binary-format strings

get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 20:45:28 +02:00
Niels Lohmann
515d994acb 📄 adjust year (#5044)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-01-01 20:00:39 +01:00
Niels Lohmann
54be9b04f0 📄 update REUSE (#4960) 2025-10-23 06:56:36 +02:00
Niels Lohmann
1705bfe914 🔖 set version to 3.12.0 (#4727)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2025-04-11 10:41:14 +02:00
Niels Lohmann
f06604fce0 Bump the copyright years (#4606)
* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
2025-01-19 17:04:17 +01:00
Niels Lohmann
620034ecec ♻️ allow patch and diff to be used with arbitrary string types (#4536) 2024-12-13 07:24:50 +01:00