Commit Graph

1092 Commits

Author SHA1 Message Date
Niels Lohmann
b6f0daf319 Write compact dumps of json_view without a library call per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer: the same walk and output, with the write position in a local
variable (stores through char pointers would otherwise force a reload
of the buffer's members after each one), strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left (the buffer keeps 64 bytes of slack), and decoded
strings copied in runs up to the next quote, backslash, or control
character. Documents that are not edited are walked through the node
array in order, so that a frame only needs the end of its container,
and integer tokens are read from the source directly. The innermost
open container is kept in local variables, and the stack holds only
the ones around it; the stack starts in a local array of 32 and moves
to the heap only for deeper nesting (its address does not escape, so
its pointers stay in registers). Dumps of shallow documents thus
allocate only the output, whose first size includes the slack, so it
does not grow just before the end.

The long copies are out of line: otherwise, the compiler merges the
fixed-size moves into the same library call.

These techniques come from the prototype; the writer lost them when the
view was split into pull requests, which made dump() 2 to 3 times
slower.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:23 +02:00
Niels Lohmann
f84524e59a Address the cpplint findings of json_document images
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:19 +02:00
Niels Lohmann
ec85028de0 Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).

- save() writes an edited document in its current state, in document
  order (floats that are not finite become null, as in dump()); the
  same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
  load(vector&&) keeps it without a copy. The nodes are copied (aligned,
  and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
  bounds, UTF-8, strings of the source, number tokens and their values);
  bounds checks structure and bounds, so that reading and serializing
  stay safe; none trusts the image.

A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.

As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.

The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.

Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:16 +02:00
Niels Lohmann
772ccbea34 Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
  number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names

The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.

Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:10 +02:00
Niels Lohmann
ee5cb0a847 Cover the edit storage in tests
Test edits of an empty document and assignments through a view of a
value that is no longer part of the document; copy the entries of a
block through deref() (one path for links and values); mark the
4 GiB limit and the returns of find_parent() that no document reaches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:06 +02:00
Niels Lohmann
e7bf81f040 Address the clang-tidy findings of the editable documents
Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:01 +02:00
Niels Lohmann
fd3a2db9ab Add editable documents with set() and push_back()
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:

- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
  object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
  the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)

Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.

Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.

Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:59 +02:00
Niels Lohmann
3806caebd8 Walk the index of json_view through a navigation policy
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:

- node flags `edited` (a string or number token in the edit arena),
  `moved` (the elements of an array/object live in a separate sequence),
  and `is_new` (no source position), and link nodes (kind_link) that
  stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
  values, and the edit arena

materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:57 +02:00
Niels Lohmann
0f6438a469 Fix old clang: do not declare the defaulted document_data() noexcept
With the nested struct object_index, clang 4 (and, by the same bug, the
clang 3.x of ci_test_compilers_clang) rejects the explicitly noexcept
defaulted constructor: "default member initializer for 'indexes' needed
within definition of enclosing class 'document_data' outside of member
functions". Nothing depends on the constructor being noexcept, so let it
take the implicit exception specification.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:54 +02:00
Niels Lohmann
cb69a5fbce Fix CI: useless casts of the key hash of the view's object index
GCC -Werror=useless-cast on Linux x86-64 rejects
static_cast<std::size_t>(key_hash(...)): the call returns a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Store the hash in a variable
and cast that, which GCC does not report.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:52 +02:00
Niels Lohmann
c05764cbda Index the large objects of json_view
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.

Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.

Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:48 +02:00
Niels Lohmann
d33586ca34 Address the clang-tidy findings of the SIMD scan
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:45 +02:00
Niels Lohmann
eb6c92e899 Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:43 +02:00
Niels Lohmann
cbd93ff4bc Convert the floats of json_view from the digit layout
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For floats and doubles with at most 19
digits, the value is now read from that layout: the digits eight at a time,
without scanning the token, and rounded by the library's conversion core
(detail::decimal_to_float(): Clinger's fast path where both operands are
exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19
digits). It rounds correctly, so the values are those of parse(); other
tokens and types keep the library's conversion of the whole token.

get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s.

Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and
those of float) to the bit-for-bit comparison with parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:37 +02:00
Niels Lohmann
a956724170 Address the cpplint findings of json_view's comparisons
compare.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:34 +02:00
Niels Lohmann
d40d049c14 Address the clang-tidy findings of the comparisons
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:32 +02:00
Niels Lohmann
af101d5e5b Add comparisons to json_view
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.

Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:28 +02:00
Niels Lohmann
5199259a63 Mark the cases of the view's serializer that tests cannot reach
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:25 +02:00
Niels Lohmann
71b47dd1de Address the clang-tidy findings of dump()
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:23 +02:00
Niels Lohmann
82cabbdc6b Add dump() to json_view
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.

The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.

Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:19 +02:00
Niels Lohmann
cf4bd21c81 Fix CI: useless cast in the array index check of the view's JSON pointers
GCC -Werror=useless-cast on Linux x86-64 rejected
static_cast<std::uint64_t>((std::numeric_limits<std::size_t>::max)()),
as both are the same type there. Compare without the cast: std::size_t
converts to std::uint64_t implicitly on every platform.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:17 +02:00
Niels Lohmann
7767ce246d Inline get() of arithmetic values of json_view
get<T>() of arithmetic types is inlined down to the conversion, so that
its checks of the node kind merge with those of the caller, and reading
an integer needs no call. Traversing every value: citm_catalog -6%,
marine_ik -5%, numbers and twitter -3%, mesh -2.5%, canada -1% (and
more above the float conversion from the digit layout: citm_catalog
-14%, marine_ik -11%).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:13 +02:00
Niels Lohmann
2161794b5a Address the clang-tidy findings of values and JSON pointers
get_string() and number_token() return braced lists; the test compares
floats by their bit patterns instead of with memcmp, uses std::any_of, and
marks a fixed seed, a default member initializer (needed by GCC's
-Weffc++), and a string search.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:10 +02:00
Niels Lohmann
8b58c8cd7c Add values and JSON pointers to json_view
basic_json_view gains get<T>(), get_to(), value() with keys and JSON
pointers, and operator[], at(), and contains() with JSON pointers, plus
two functions basic_json has no counterpart for:

- get_string(): the string without a copy (a string_view into the source,
  or into the decoded strings for strings with escapes)
- number_token(): the text of a number as it appears in the source

get<T>() converts arithmetic types, strings (also string_view_t),
std::nullptr_t, std::vector, maps with string keys, and views directly;
floats are converted from the digit layout recorded by the parser with the
library's conversion chain, so the values are bit-identical to parse().
Other types, including user types with from_json(), go through
materialize().

The exceptions are those of basic_json, message included. Where const
basic_json has undefined behavior (a missing key or an index out of range
with operator[] and a JSON pointer), the result is a discarded view;
value() returns the default wherever basic_json catches out_of_range, and
contains() never throws. Array indices of JSON pointers follow
json_pointer's rules (parse_error.106/109, out_of_range.404/410).

Tests compare the conversions of 2,000 generated documents, 20,000 float
tokens (double and float, bit for bit), and every JSON pointer of 1,000
documents with basic_json, and the exceptions for malformed pointers.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:07 +02:00
Niels Lohmann
a7ad17d274 Give code outside basic_json the reference tokens of a json_pointer
detail::json_pointer_access returns the reference tokens of a pointer, so
that code resolving pointers without a basic_json value (such as the
zero-copy view) does not have to parse to_string() again. No change in
behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:05 +02:00
Niels Lohmann
673e3c881a Address the clang-tidy findings of element access and iteration
Marks the default initializer of the item's index string (needed by GCC's
-Weffc++) and, in the test, an escaped literal and a comparison of find()
with end(), which is what the test is about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:01 +02:00
Niels Lohmann
5c812a5507 Add element access, lookup, and iteration to json_view
basic_json_view gains the read-only access functions of basic_json:
operator[] and at() with keys and indices, front(), back(), find(),
contains(), count(), begin()/end(), items() (with structured bindings from
C++17 on), and type_name(). They throw the exceptions (ids and messages)
that the const functions of basic_json throw; where basic_json has
undefined behavior (operator[] with a missing key or an index out of
range, front()/back() of an empty container), the view returns a
discarded view or throws invalid_iterator.214.

Objects are iterated in document order, and all members are visited. With
duplicate keys, lookups find the first member, so that a lookup can stop
at the first match; parse() keeps the last value. Keys of up to 16 bytes
are compared with two overlapping loads instead of memcmp, and most keys
are rejected by their length alone, from the index.

The iterators and items live in detail/view/iterator.hpp, the lookups in
detail/view/lookup.hpp. Tests compare every element and member of 2,000
generated documents with ordered_json, keys of every length around the
load sizes, the exceptions against const basic_json, and the iterators.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:58 +02:00
Niels Lohmann
6cda4b8461 Name JSON types without a basic_json value
basic_json::type_name() now calls detail::value_type_name(value_t), so
that code which reports types without a basic_json value at hand, such as
the zero-copy view, uses the same names in its exception messages. No
change in behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:56 +02:00
Niels Lohmann
ab2cf47785 Address the cpplint findings of json_view's materialize()
materialize.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:28 +02:00
Niels Lohmann
1d0d28c72f Mark the code of json_view that tests cannot reach
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:25 +02:00
Niels Lohmann
b1628423a6 Address the clang-tidy findings of json_document and json_view
- the input dispatch takes byte ranges by const reference and reads the
  size once (which also settles a finding of the static analyzer); input
  adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
  constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
  the other tests

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:23 +02:00
Niels Lohmann
34f1bff147 Add json_document and json_view: parse, accept, types, materialize
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:

- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
  inputs, owning rvalue strings, streams, and other inputs), parse_copy,
  accept, read, root, is_discarded, source, owns_source, node_count,
  memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
  materialize (the value parse() would produce, built by the same SAX
  handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
  ordered_json_view

A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.

The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:04:17 +02:00
Niels Lohmann
dc9899fde6 Fix CI: useless cast in the growth of the view's node index
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:24 +02:00
Niels Lohmann
03cd4c6a95 Fix CI: MSVC C4127 in the view builder and the single-header test build
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
  constant) for `TrailingCommas && cur() == ']'` and the like when the
  option is off. Route the template arguments through a static enabled()
  function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
  single_include/, which does not contain the internal
  nlohmann/detail/view headers. Build that test only with the multiple
  headers.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:23 +02:00
Niels Lohmann
02ed84a878 Remove the eight-digit case of the view's digit parser
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:21 +02:00
Niels Lohmann
243ed0a6ff Address the cpplint findings of the view's parser
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:20 +02:00
Niels Lohmann
d821b721c1 Read whole blocks of eight digits of json_view directly
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:19 +02:00
Niels Lohmann
8f2bcb4ac9 Classify a NUL inside a string of json_view as a control character
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:01:18 +02:00
Niels Lohmann
9a20019005 Address the clang-tidy findings of the view's parser
- tables as std::array; the frames of the first 64 levels stay a C array
  (not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
  second table
- the parse failure is private, with an accessor; the special member
  functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
  branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
  the other tests

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:59:46 +02:00
Niels Lohmann
be6799a3c6 Keep the arguments of the view's throw helpers used without exceptions
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:59:45 +02:00
Niels Lohmann
23e517a64a Add the one-pass parser of json_view (internal)
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.

Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.

Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:59:44 +02:00
Niels Lohmann
e6fe273f23 Add the node index and scanning primitives of json_view (internal)
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:

- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
  its own prefix (json.hpp undefines its own at its end); the throw macro
  honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
  (checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
  arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
  (after yyjson) and the library's SWAR and UTF-8 checks, independent of the
  byte order

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:56:32 +02:00
Niels Lohmann
68baa37bc8 Keep the behavior-changing configuration readable after json.hpp
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:56:30 +02:00
Niels Lohmann
6c32ac94cf Decode \u escapes with a table in the lexer
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.

json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.

Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:56:28 +02:00
Niels Lohmann
c139eca9ba Find the stop byte of a string run without a byte loop
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.

The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:56:25 +02:00
Niels Lohmann
b0db8653b6 Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:

- The token is split into sign, significand w (at most 19 digits), and
  decimal exponent q, using the positions of the decimal point and the
  exponent that the scanners already recorded, so no character is
  classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
  an exact big-integer comparison with the midpoint between the two
  candidates (the digit comparison of fast_float, simplified).

This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.

convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.

Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.

Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 14:55:08 +02:00
Niels Lohmann
f295970fb6 Convert floats with Eisel-Lemire when std::from_chars is unavailable
Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).

The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.

json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:49:25 +02:00
Niels Lohmann
f6c115a9a6 Move the float conversion chain out of the lexer
lexer::convert_number() converted float tokens with std::from_chars (when
available), Clinger's fast path, and the locale-aware strtod fallback, all
as lexer members. They are now free functions in number_parse.hpp:

- convert_float_fast(): std::from_chars, then Clinger's fast path, skipped
  when the mantissa has too many significant digits
- convert_float_locale_aware(): strtof/strtod/strtold with the decimal point
  of the current locale, retried when the locale changed (#5198)

so that other code converting JSON number tokens gets the same values. No
change in behavior; the lexer no longer includes <clocale> and <cstdlib>.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:49:24 +02:00
Niels Lohmann
fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann
9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00