* Fix math formulas rendering as raw TeX in the published docs The privacy plugin self-hosts MathJax 2.7.0 but drops its ?config=TeX-MML-AM_CHTML query string, so the rehosted script loads no input jax and the 10 formulas across 6 pages render as raw TeX to readers. Remove pymdownx.arithmatex and the MathJax extra_javascript entry, and rewrite the formulas in plain HTML (<sup>, <i>) instead. This also drops a nine-year-old third-party script from every page. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix publish_documentation triggers and persist unneeded git credentials publish_documentation.yml only triggered on docs/mkdocs/** pushes, but the site also embeds .github/CODE_OF_CONDUCT.md, CONTRIBUTING.md, SECURITY.md, cmake/{clang,gcc}_flags.cmake, .clang-tidy, tools/astyle/.astylerc and tests/fmt_formatter/project/main.cpp via pymdownx.snippets, so changes to those files never republished the site. Extend the path filter to cover them, and switch runs-on from the long-pinned ubuntu-22.04 to ubuntu-latest to match ci_test_documentation. Also add persist-credentials: false to the checkouts in ci_icpx and ci_nvhpc (ubuntu.yml) and msvc-vs2026/msvc-arm64 (windows.yml), none of which pushes with git, matching every other checkout in these workflows. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix example build: broken debug echo, deprecations hidden for everything docs/Makefile's debug echo used a space instead of a comma in $(call cxx_standard ...), so it always printed an empty standard. Every example was also compiled with -Wno-deprecated-declarations, which would silently hide an accidental deprecated-API call in any of them. Factor the duplicated compile flags into EXAMPLE_CPPFLAGS/ EXAMPLE_WARNFLAGS, build with -Werror=deprecated-declarations by default, and only allow the three examples that intentionally document deprecated API (the DEPRECATED_EXAMPLES list) to suppress it. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix check_structure.py NOLINT parsing and an off-by-one report line `line.strip("<!-- NOLINT")` followed by `.strip(" -->")` strips any of the characters in those sets from both ends, not a literal prefix; it only happened to work for "Examples". A NOLINT'd section name starting with N, O, L, I or T (e.g. "Notes", "Template parameters", "Iterator invalidation", "Literals") was silently mangled, so the suppression did not apply and the checker could report a spurious missing/misordered section. Parse the comment with a regex instead. The same fragile strip() pattern was used for heading text; replace it with a plain prefix slice. Also fix the admonition_title report, which used the 0-based line index while every other report in the file uses lineno+1. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix docset README's non-existent make target and stale fallback URL The README told readers to run `make nlohmann_json.docset`, but the Makefile's targets are `all`, `JSON_for_Modern_C++.docset` and `install_docset_zeal`; the documented command has failed with "No rule to make target" since #2967 (2021). Point the README at the real target and folder name. Info.plist's DashDocSetFallbackURL also still pointed at the old nlohmann.github.io/json/ URL instead of the canonical https://json.nlohmann.me/ from mkdocs.yml's site_url. Leave list_missing_pages/list_removed_paths alone: they may become redundant once #5638's check_docset() lands, which is a follow-up. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove two leftover Doxygen-era .link files from the examples directory parse__iterator_pair.link and parse__pointers.link each held only a Wandbox "online" permalink from the old Doxygen docs. #3071 deleted every other .link file in 2021; these two came in through a parallel PR (#3100) and were never referenced by any page, script or config. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix MacPorts CMake example to include its own snippet files The MacPorts "Example: CMake" block included integration/homebrew/example.cpp and integration/homebrew/CMakeLists.txt instead of the MacPorts files right next to it, a copy-paste slip from the Homebrew section. Nothing referenced integration/macports/CMakeLists.txt as a result. The page rendered correctly only because the homebrew, macports and vcpkg/CMakeLists.txt snippets are byte-identical, so a future edit to the MacPorts files would not have shown up on the page. Part of #5718 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Do not persist git credentials in publish_documentation's checkout The checkout step in publish_documentation.yml left the default persist-credentials: true, so GITHUB_TOKEN stayed writable in .git/config for the rest of the job (zizmor's artipacked finding). The Deploy documentation step authenticates through its own github_token input to peaceiris/actions-gh-pages and does not push with the checked-out credentials, so persist-credentials: false is safe here, matching every other checkout in the workflow set. Overlaps #5638, which edits this same checkout step (adds fetch-depth: 0); expect a rebase conflict there. Part of #5718 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Replace list_missing_pages/list_removed_paths with a comm(1)-based diff The docset Makefile's list_missing_pages ran one sqlite3 query per mkdocs page, and list_removed_paths nested a loop over all mkdocs pages inside a loop over all docset index paths (O(n*m) shell iteration). Issue #5718 item 5 suggested removing or reducing these targets once #5638's check_docset() lands, but that PR is still open and covers only API pages and macros, not the full page set these targets check. Replace the loops with two sorted path lists (DOCSET_PAGE_PATHS from mkdocs' markdown sources, DOCSET_INDEX_PATHS from the built docset index) compared with a single comm(1) call each, verified to produce output identical to the old loops against the current docSet.dsidx. The sed expression used '#' as its delimiter, which GNU Make reads as a comment character even inside a variable assignment, truncating the line and orphaning the closing paren of $(shell ...) ("unterminated call to function 'shell': missing ')'"). Use '@' as the delimiter instead. Part of #5718 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
13 KiB
BJData
The BJData format was derived from and improved upon
Universal Binary JSON(UBJSON) specification (Draft 12). Specifically, it introduces an optimized
array container for efficient storage of N-dimensional packed arrays (ND-arrays); it also adds 5 new type markers -
[u] - uint16, [m] - uint32, [M] - uint64, [h] - float16 and [B] - byte - to unambiguously map common binary
numeric types; furthermore, it uses little-endian (LE) to store all numerics instead of big-endian (BE) as in UBJSON to
avoid unnecessary conversions on commonly available platforms.
Compared to other binary JSON-like formats such as MessagePack and CBOR, both BJData and UBJSON demonstrate a rare combination of being both binary and quasi-human-readable. This is because all semantic elements in BJData and UBJSON, including the data-type markers and name/string types, are directly human-readable. Data stored in the BJData/UBJSON format is not only compact in size, fast to read/write, but also can be directly searched or read using simple processing.
!!! abstract "References"
- [BJData Specification](https://neurojson.org/bjdata/draft2)
Serialization
The library uses the following mapping from JSON values types to BJData types according to the BJData specification:
| JSON value type | value/range | BJData type | marker |
|---|---|---|---|
| null | null |
null | Z |
| boolean | true |
true | T |
| boolean | false |
false | F |
| number_integer | -9223372036854775808..-2147483649 | int64 | L |
| number_integer | -2147483648..-32769 | int32 | l |
| number_integer | -32768..-129 | int16 | I |
| number_integer | -128..127 | int8 | i |
| number_integer | 128..255 | uint8 | U |
| number_integer | 256..32767 | int16 | I |
| number_integer | 32768..65535 | uint16 | u |
| number_integer | 65536..2147483647 | int32 | l |
| number_integer | 2147483648..4294967295 | uint32 | m |
| number_integer | 4294967296..9223372036854775807 | int64 | L |
| number_integer | 9223372036854775808..18446744073709551615 | uint64 | M |
| number_unsigned | 0..127 | int8 | i |
| number_unsigned | 128..255 | uint8 | U |
| number_unsigned | 256..32767 | int16 | I |
| number_unsigned | 32768..65535 | uint16 | u |
| number_unsigned | 65536..2147483647 | int32 | l |
| number_unsigned | 2147483648..4294967295 | uint32 | m |
| number_unsigned | 4294967296..9223372036854775807 | int64 | L |
| number_unsigned | 9223372036854775808..18446744073709551615 | uint64 | M |
| number_float | any value | float64 | D |
| string | with shortest length indicator | string | S |
| array | see notes on optimized format/ND-array | array | [ |
| object | see notes on optimized format | map | { |
| binary | see notes on binary values | array | [$B |
!!! success "Complete mapping"
The mapping is **complete** in the sense that any JSON value type can be converted to a BJData value.
Any BJData output created by `to_bjdata` can be successfully parsed by `from_bjdata`.
!!! warning "Size constraints"
The following values can **not** be converted to a BJData value:
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! info "Unused BJData markers"
The following markers are not used in the conversion:
- `Z`: no-op values are not created.
- `C`: single-byte strings are serialized with `S` markers.
!!! info "NaN/infinity handling"
If NaN or Infinity are stored inside a JSON number, they are serialized properly. This behavior differs from the
`dump()` function which serializes NaN or Infinity to `#!json null`.
!!! info "Endianness"
A breaking difference between BJData and UBJSON is the endianness of numerical values. In BJData, all numerical data
types (integers `UiuImlML` and floating-point values `hdD`) are stored in the little-endian (LE) byte order as
opposed to big-endian as used by UBJSON. Adopting LE to store numeric records avoids unnecessary byte swapping on
most modern computers where LE is used as the default byte order.
!!! info "Optimized formats"
Optimized formats for containers are supported via two parameters of
[`to_bjdata`](../../api/basic_json/to_bjdata.md):
- Parameter `use_size` adds size information to the beginning of a container and removes the closing marker.
- Parameter `use_type` further checks whether all elements of a container have the same type and adds the type
marker to the beginning of the container. The `use_type` parameter must only be used together with
`use_size = true`.
Note that `use_size = true` alone may result in larger representations - the benefit of this parameter is that the
receiving side is immediately informed of the number of elements in the container.
!!! info "ND-array optimized format"
BJData extends UBJSON's optimized array **size** marker to support ND-arrays of uniform numerical data types
(referred to as *packed arrays*). For example, the 2-D `uint8` integer array `[[1,2],[3,4],[5,6]]`, stored as nested
optimized array in UBJSON `[ [$U#i2 1 2 [$U#i2 3 4 [$U#i2 5 6 ]`, can be further compressed in BJData to
`[$U#[$i#i2 2 3 1 2 3 4 5 6` or `[$U#[i2 i3] 1 2 3 4 5 6`.
To maintain type and size information, ND-arrays are converted to JSON objects following the **annotated array
format** (defined in the [JData specification (Draft 3)][JDataAAFmt]), when parsed using
[`from_bjdata`](../../api/basic_json/from_bjdata.md). For example, the above 2-D `uint8` array can be parsed and
accessed as
```json
{
"_ArrayType_": "uint8",
"_ArraySize_": [2,3],
"_ArrayData_": [1,2,3,4,5,6]
}
```
Likewise, when a JSON object in the above form is serialized using
[`to_bjdata`](../../api/basic_json/to_bjdata.md), it is automatically converted into a compact BJData ND-array.
When parsing, an ND-array whose dimension vector is empty, contains a single integer, contains two integers with the
first being 1, or contains a 0 is returned as a regular (possibly empty) array rather than an annotated object.
An object is only converted if the annotation describes a packed array that is parsed back into the same annotated
object; otherwise it is serialized as a regular JSON object, so the annotation is never lost in a round trip. This requires
all of the following:
- `"_ArrayType_"` is one of `uint8`, `int8`, `uint16`, `int16`, `uint32`, `int32`, `uint64`, `int64`, `single`,
`double`, `char`, or `byte`,
- `"_ArraySize_"` is an array, since the dimensions are written as the ND-array header's length,
- `"_ArraySize_"` has at least two entries and is not a 1×N row vector (first entry 1), since other shapes are
parsed back as a regular array,
- every entry of `"_ArraySize_"` is a positive integer, and their product is representable as a `std::size_t`,
- `"_ArrayData_"` is an array holding exactly that many elements, and
- every element of `"_ArrayData_"` is a number of the kind named by `"_ArrayType_"` (a floating-point number for
`single` and `double`, an integer otherwise).
The current version of this library does not yet support automatic detection of and conversion from a nested JSON
array input to a BJData ND-array.
[JDataAAFmt]: https://github.com/NeuroJSON/jdata/blob/master/JData_specification.md#annotated-storage-of-n-d-arrays
!!! info "Restrictions in optimized data types for arrays and objects"
Due to diminished space saving, hampered readability, and increased security risks, in BJData, the allowed data
types following the `$` marker in an optimized array and object container are restricted to
**non-zero-fixed-length** data types. Therefore, the valid optimized type markers can only be one of
`UiuImlMLhdDCB`. This also means other variable (`[{SH`) or zero-length types (`TFN`) can not be used in an
optimized array or object in BJData.
!!! info "Binary values"
BJData provides a dedicated `B` marker (defined in the [BJData specification (Draft 3)][BJDataBinArr]) that is used
in optimized arrays to designate binary data. This means that, unlike UBJSON, binary data can be both serialized and
deserialized.
To preserve compatibility with BJData Draft 2, the Draft 3 optimized binary array must be explicitly enabled using
the `version` parameter of [`to_bjdata`](../../api/basic_json/to_bjdata.md).
In Draft2 mode (default), if the JSON data contains the binary type, the value stored as a list of integers, as
suggested by the BJData documentation. In particular, this means that the serialization and the deserialization of
JSON containing binary values into BJData and back will result in a different JSON object.
[BJDataBinArr]: https://github.com/NeuroJSON/bjdata/blob/master/Binary_JData_Specification.md#optimized-binary-array
??? example
```cpp
--8<-- "examples/to_bjdata.cpp"
```
Output:
```c
--8<-- "examples/to_bjdata.output"
```
Deserialization
The library maps BJData types to JSON value types as follows:
| BJData type | JSON value type | marker |
|---|---|---|
| no-op | no value, next value is read | N |
| null | null |
Z |
| false | false |
F |
| true | true |
T |
| float16 | number_float | h |
| float32 | number_float | d |
| float64 | number_float | D |
| uint8 | number_unsigned | U |
| int8 | number_integer | i |
| uint16 | number_unsigned | u |
| int16 | number_integer | I |
| uint32 | number_unsigned | m |
| int32 | number_integer | l |
| uint64 | number_unsigned | M |
| int64 | number_integer | L |
| byte | number_unsigned | B |
| string | string | S |
| char | string | C |
| array | array (optimized values are supported) | [ |
| ND-array | object (in JData annotated array format) | [$.#[. |
| object | object (optimized values are supported) | { |
| binary | binary (strongly-typed byte array) | [$B |
!!! success "Complete mapping"
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
[`to_bjdata`](../../api/basic_json/to_bjdata.md) using any combination of options and parsed back into an equal
value, and serializing that value again with the same options produces the same bytes. The exception is binary
values: they are only written as an optimized binary array (`[$B`) if Draft 3 is enabled and both `use_size` and
`use_type` are set. Otherwise, they are written as arrays of integers and parsed back as such (see the notes on
binary values above), and serializing such an array again may choose different, but equally valid, type markers.
The bytes can then differ, but parsing them again yields the same value.
??? example
```cpp
--8<-- "examples/from_bjdata.cpp"
```
Output:
```json
--8<-- "examples/from_bjdata.output"
```