Commit Graph

17 Commits

Author SHA1 Message Date
Niels Lohmann
ec120e78c8 Document editable json_documents
- API pages for set and push_back of basic_json_document and for the
  editable aliases; the class overviews name the Editable parameter
- type_error.319 on the exceptions page
- the feature page explains editing, and the examples show when it
  helps: a configuration file changed without reformatting it

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:09:04 +02:00
Niels Lohmann
eb6c92e899 Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:43 +02:00
Niels Lohmann
76a0e3befd Document the comparisons of json_view
- API pages for operator== and operator!= of basic_json_view, linked
  both ways with the basic_json pages
- the feature page and the class overview list the comparisons
- the examples show when the view helps: detecting a changed document
  without building json values

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:31 +02:00
Niels Lohmann
f1edbabd3e Document dump() of json_view
- API pages for dump, number_format, and operator<< of basic_json_view,
  linked both ways with the basic_json pages
- the feature page describes document order and number_format::source
- the examples show when the view helps: forwarding part of a message
  and writing numbers exactly as they were read

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:22 +02:00
Niels Lohmann
d3d2ad3e4a Document values and JSON pointers of json_view
- API pages for get, get_to, get_string, number_token, and value of
  basic_json_view; JSON pointer overloads of operator[], at, and
  contains; links both ways with the basic_json pages
- the feature page describes which conversions copy nothing
- the examples show when the view helps: strings without copies, numbers
  exactly as written, and paths into a large text

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:09 +02:00
Niels Lohmann
2f087c6b73 Document element access, lookup, and iteration of json_view
- API pages for operator[], at, front, back, find, contains, count,
  begin, end, cbegin, cend, items, and type_name of basic_json_view,
  linked both ways with the basic_json pages
- the feature page and size() describe document order and duplicate
  keys
- the examples show when the view helps: reading a few fields of a large
  text, probing optional members, and members in source order

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:08:00 +02:00
Niels Lohmann
2d3492dccb Document json_document and json_view
- API pages for basic_json_document and basic_json_view, one per member,
  and for the four aliases, each with an example
- features/json_view.md: the problem the view solves, ownership and
  lifetime, what matches basic_json::parse() and what differs, and when
  to choose json, ordered_json, SAX, or the view
- the examples show why one would use the view, not only how: borrowed
  vs. owned input, reading a few fields and materializing one subtree,
  reusing a document across many messages
- registered in the mkdocs navigation, llms.txt, the docset, the
  exceptions page (out_of_range.416), architecture.md, the integration
  page, and the README; the yyjson credit is added to the README and
  license.md

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 15:07:21 +02:00
Niels Lohmann
1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00
Niels Lohmann
26cfec34be Clean up and document project files (#4560) 2025-01-17 06:53:35 +01:00
Niels Lohmann
885aa0014e Add migration guide (#3887) 2022-12-18 17:04:23 +01:00
Florian Albrechtskirchinger
0e61ee8b07 Restructure inline namespace and allow version component to be disabled (#3683) 2022-08-10 20:24:16 +02:00
Florian Albrechtskirchinger
d79d12f854 Update docset index (#3686) 2022-08-07 20:33:15 +02:00
Florian Albrechtskirchinger
9aafcbe965 Move UDLs out of the global namespace (#3605)
* Move UDLs into nlohmann::literals::json_literals namespace

* Add 'using namespace' to unit tests

* Add 'using namespace' to examples

* Add 'using namespace' to README

* Move UDL mkdocs pages out of basic_json/

* Update documentation

* Update docset index

* Add JSON_GlobalUDLs CMake option

* Add unit test

* Build examples without global UDLs

* Add CI target
2022-07-31 17:38:52 +02:00
Florian Albrechtskirchinger
d3e347bd2d More documentation updates for 3.11.0 (#3553)
* mkdocs: add string_view examples

* mkdocs: reference underlying operators

* mkdocs: add operator<=> examples

* mkdocs: fix style check issues

* mkdocs: tweak BJData page

* mkdocs: add CMake option hints to macros

* mkdocs: fix JSON_DISABLE_ENUM_SERIALIZATION definition

* mkdocs: fix link to unit-udt.cpp

* mkdocs: fix "Arbitrary Type Conversions" title

* mkdocs: link to api/macros/*.md instead of features/macros.md

* mkdocs: document JSON_DisableEnumSerialization CMake option

* mkdocs: encode required C++ standard in example files

* docset: detect gsed/sed

* docset: update index

* docset: fix CSS patching

* docset: add list_missing_pages make target

* docset: add list_removed_paths make target

* docset: replace page titles with name from index

* docset: add install target for Zeal docset browser

* Use GCC_TOOL in ci_test_documentation target
2022-07-31 14:05:58 +02:00
Niels Lohmann
b185c5bc9d 📝 clean up after #3581 (#3596) 2022-07-22 15:35:29 +02:00
Richard Hozák
f6acdbec2c Allow disabling default enum conversions (#3536) 2022-06-16 19:34:32 +02:00
Niels Lohmann
b21c345179 Reorganize directories (#3462)
* 🚚 move files
* 🚚 rename doc folder to docs
* 🚚 rename test folder to tests
2022-05-01 09:41:50 +02:00