Skip to content

Format Stability Policy

Every document and every store in this contract carries an explicit format and format_version pair. Read format_version before anything else — this page is what those numbers promise, so a client can detect a future change in code instead of assuming today’s shape is permanent.

The index document’s format/format_version ("nemar-zarr-index" / 3) and the store’s own root format/format_version ("biosigio-zarr" / 2) are separate counters that travel independently. A future index format change does not imply a store format change, or the reverse — check whichever one the field you are reading lives in.

shared/zarr-index.schema.json (nemarOrg/nemar-cli, served at GET /schemas/zarr-index-v3.json) sets additionalProperties: false on every object in the document. This makes format version 3 closed: the converter validates every index against this exact schema before publishing (validate_document in scripts/zarr/generate_zarr.py), and refuses to publish a document that does not conform — so an index this schema accepts is, by construction, an index containing nothing this page does not already describe.

  • Additive within a version. A change that only adds new optional fields ships as the same format_version (3), with the new fields declared in the schema. A client pinned to v3 and ignoring fields it does not recognize keeps working unchanged.
  • Breaking is a new version, at a new path. A removed or retyped field, or a narrowed enum, ships as format_version 4, served at a new schema path (/schemas/zarr-index-v4.json) rather than edited in place at v3’s path. /schemas/zarr-index-v3.json keeps serving the v3 schema for as long as any dataset still publishes a v3 index — there is no fixed sunset date, because the index is the mandatory entry point (anonymous ListBucket is denied; see Access and hosting) and a client has nothing else to fall back to if it disappeared.

The store: additive so far, on two different floors

Section titled “The store: additive so far, on two different floors”

biosigIO’s own store format_version has stayed at 2 through several rounds of new attributes, by the same rule the index uses: a reader that ignores attributes it does not recognize keeps working. Three of those rounds are not on the same footing, though:

  • The sss root attribute (Signal-Space Separation disclosure, ADR 0028) is already live in production today — it predates the v3 rollout.
  • The declared pyramid and chunk-geometry attributes (n_view_levels, view_levels, chunk_seconds, shard_seconds, chunk_samples, shard_samples, source_rate_hz, view_chunk_columns) need biosigIO ≥1.2.6 and are present in current converted stores.
  • The nemar root attribute is written by the NEMAR converter itself, not by biosigIO, and is present in current v3 conversions regardless of the biosigIO version that supplies the other fields.

See the store contract’s current-deployment note for the biosigIO version split behind channels_tsv_units/bids_unit specifically, and for what a live production store looks like today (checked 2026-09-09: biosigio_version: "1.2.7", with the current optional fields present on the sampled store).

biosigIO’s stated policy is that a reader should reject a store whose format_version is newer than the one it supports, rather than guess at an unfamiliar layout — the same “read the version, don’t assume the shape” discipline the index asks for.

Because the store’s own version has not moved, there is no store-side deprecation window to describe: current stores remain biosigIO format version 2. The store contract page’s current-deployment note is about which optional attributes a given store happens to carry, tied to the biosigIO and producer versions that wrote it, not about a store format bump.

  • The schema itself. GET /schemas/zarr-index-v3.json is the canonical, machine-readable description of the index; a schema validator failing against a live document is the most direct signal something changed.
  • ADR 0033 governs when a producer-side change reaches an already-converted dataset (only on engine_version bump — never assume a change is live everywhere just because it merged).
  • This page’s own history, in nemarOrg/docs — a format bump here is a documentation change reviewed the same way any other pull request is.
  • nemarOrg/nemar-cli releases — the index and store schemas live in that repository (shared/*.schema.json), and a format bump ships as part of a normal release, following the release pipeline documented there.

Everything promised above is re-checked against every index the catalog publishes, not sampled and not asserted once at conversion time. The checker verifies footer entry-count divisibility, that each index agrees with the array it describes, that inner chunks span every channel, the codec chain and dtype, per-channel scale/offset, and that every events.parquet column is ZSTD-compressed.

This exists because a measured-then-trusted claim already failed once: a shard-index rule was measured on two stores, written into a design document, and then believed by the code, the tests, the fixture and five reviewers alike, which shipped a reader that returned signal from the wrong place in the recording. A claim about on-disk geometry is only as good as the command that re-checks it.

Maintainers run it on demand rather than per pull request; the procedure is in the operations runbook.

How verification works, and what it will not claim

Section titled “How verification works, and what it will not claim”

A store is verified by a standing sweep, not only by a gate at conversion time. Two filters name the same dataset differently and the difference is deliberate:

  • has_zarr means converted: zarr_status = 'ready' with at least one store. It has kept that meaning since it was introduced, so a caller filtering on it never sees a silently narrower result set.
  • has_zarr_verified is the stricter one: converted and the sweep’s last verdict was verified.

A freshly converted dataset reads zarr_verify_status: null until the daily sweep reaches it. Serving never waits for verification (ADR 0005): verification is reported, never a precondition for reading data.

These are worth stating publicly rather than leaving in one repository’s internal notes, because they explain results you can observe.

It re-derives ground truth from the dataset’s own BIDS metadata over the public, credential-free raw.githubusercontent.com content host, and never the GitHub API, App, or a personal token. That is what makes it safe to run outside production, and it is why candidates are restricted to status='active' AND visibility='public': a private repository cannot be read anonymously, so including one would only ever manufacture unverifiable noise.

It fails open on the row, never on the verdict. A transient error, a non-2xx that is not a 404, a network failure, or either fetch budget running out mid-dataset aborts that one dataset’s verification for the run. Nothing is stamped, it lands in the run’s errors, and the row stays a candidate. Only a clean 404 at every candidate path is real absence, and only a value that actually parsed counts as checked. An infrastructure problem must never be recorded as a verdict about the data.

A verdict expires when the bytes change. A re-conversion (a changed zarr_source_commit) or a null stamp re-arms verification, so a stale verdict cannot outlive the store it described.

Why a recording can sit in pending rather than failures

Section titled “Why a recording can sit in pending rather than failures”

failures[] is what will not convert without a change to the data or the converter. pending[] is the other half: recordings with no store that are still expected to convert.

Only the attempted pendings (infra_failure, memory_budget) advance the backoff — 1 hour, 6 hours, 24 hours, then weekly — and the five-round exhaustion cap, after which the recording becomes a typed retry_exhausted failure. A not_attempted pending is re-queued at the shortest delay and does not advance the retry round, because nothing has failed for it. A dataset merely too large to finish in one run must not burn its rounds on recordings nobody has tried yet, which would mass-promote them into published failures.

Units come from the recording’s own sidecar

Section titled “Units come from the recording’s own sidecar”

Served samples carry the unit the recording’s channels.tsv declares, on both conversion paths. The sidecar is resolved by BIDS inheritance and passed explicitly, never as the exporter’s "auto" detection: the file handed to the exporter is a scratch materialisation, and on the MaxShield path it is a filtered copy, so sibling detection would be asking the wrong question. units_report on a store entry is present exactly when a sidecar was applied.