# Dataset API and releases

The dataset routes are read-only. The [OpenAPI document](../apps/site/page/data/openapi.json)
also describes argument review and transcription, which use POST. Data routes
use GET/HEAD; OPTIONS supports browser preflight.

| Dataset | Route | Scope |
| --- | --- | --- |
| Lexicon | `/api/lexicon`, `/api/lexicon/words/{word}` | Definitions, examples and directed meaning links; imported source coverage |
| Authored usage | `/api/usage`, `/api/usage/contrasts` | 36 contrasts, explicitly unreviewed teaching material |
| Literary context | `/api/datasets/literary-context` | 5,096 excerpts with per-work provenance |
| Association practice | `/api/datasets/association-practice` | 1,050 prompts; evidence categories, not a validated semantic graph |
| Observed language | `/api/datasets/observed-language` | Reviewed, rights-cleared source/claim records; scoped key and private storage required |
| Argument review | `/api/argument`, `/api/argument/review` | Capabilities; local wording checks or optional source-grounded AI review |

## Argument service

`POST /api/argument/review` accepts `text` (1–12,000 characters, at most 100
sentences), optional `context` (2,000) and `reply` (12,000), and `sources`
(up to four `{id, label, text, url?}` records). Each source is bounded to
24,000 characters and 600 sentences. IDs must be unique. URLs are metadata,
never fetched. `mode` defaults to `local`. The response's `method` distinguishes
`text-checks` from `model-assisted`; neither performs external verification.
Every passage offset is a half-open UTF-16 range in the original submitted text.

Local checks are deterministic vocabulary comparisons, not entailment or
fact-checking. Related wording and heuristic cautions invite inspection.
An absent match is not a negative factual judgment. App checks run locally;
the POST route executes the same module without persisting requests.

Cloud mode requires `mode: "cloud"`, `consent: true`, and a Bearer key with
the `argument-review` entitlement in `DATA_KEYS`. Configure server-only
`OPENAI_API_KEY`, an explicit Responses-compatible `ARGUMENT_MODEL` supporting
structured output, and `ARGUMENT_RATE_LIMITER` with a `limit({key})` binding.
R2 is not required. Keys use the same SHA-256 lookup and expiry contract as
the dataset service. Missing configuration keeps cloud features unavailable.
Use a [Cloudflare rate limit binding](https://developers.cloudflare.com/workers/runtime-apis/bindings/rate-limit/)
for short bursts (for example, 5 requests / 60 seconds). This limiter is local
to each Cloudflare location and eventually consistent; it is not a global
billing quota. Set provider spend controls and measure actual costs before
offering paid allowances. Confirm binding support on the deployment target.

The provider receives supplied text plus at most eight dictionary terms
(up to eight candidate senses each), source attribution and relevant editorial
notes. Its review covers up to 12 central claims, not an exhaustive proof.
Code checks verbatim claim/source quotations and requires citations for
supported/contradicted judgments. A valid quotation does not guarantee a valid
interpretation. Malformed, refused, incomplete or untraceable reviews return
502. Requests use a 45-second provider timeout, bounded bodies, no retries and
`store: false`. This is not a zero-retention guarantee; see
[OpenAI data controls](https://developers.openai.com/api/docs/guides/your-data).

`POST /api/argument/transcribe` accepts a multipart `file` and the header
`X-Superb-Consent: send-to-provider`, with the same key and burst limit. The
entire request must fit within 20 MiB; the app permits files up to 19 MiB.
Supported extensions: MP3, MP4, M4A, WAV, WebM, MPEG, MPGA, OGG, FLAC.
`TRANSCRIPTION_MODEL` defaults to `gpt-4o-mini-transcribe`. Returned transcripts
must be reviewed before the app adds them as evidence. Visual content is
explicitly unassessed. The app imports TXT/MD/SRT/VTT locally (100 KB file
limit); PDF/DOCX must first be exported to plain text. Nothing is silently
truncated, fetched from social URLs, or collected into a dataset.

Implementation references: [structured output](https://developers.openai.com/api/docs/guides/structured-outputs),
[file transcription](https://developers.openai.com/api/docs/guides/speech-to-text).
Tests use controlled provider responses; live model quality, audio accuracy and
deployment bindings require validation against the actual configured service.

## Dataset pagination

`GET /api/datasets` discovers the routes. For the three `/api/datasets/{dataset}`
releases, append `/records` for a page, `/records/{encodedId}` for one record,
or `/download` for JSONL. Every page has `version`, `source`, `total`, `count`,
`nextCursor` and `records`. Pages hold at most 100 records. Pass `nextCursor`
unchanged as `?cursor=...`; null ends the release. Related variants are grouped
in the data and must remain together for evaluation.

Use `?version=<manifest.version>` to pin a release. A stale version or cursor
returns 409; restart from the manifest rather than merging releases. The API
serves only the current release. Old local folders are an authoring archive,
not a promise of indefinite historical hosting. Record IDs are stable within
their documented source contracts. Release versions are full SHA-256 hashes
of canonical records, source terms, schema and status.

Public data allows cross-origin access and one-hour caching. Private data and
errors use no-store. No private key, raw source capture, or unreviewed study
bundle is copied into the public assets. Missing/invalid keys return 401;
wrong dataset entitlement returns 403; unknown IDs return 404; malformed
requests return 400; missing bindings, unreviewed/revoked releases and corrupt
bytes return 503. Files are checked against manifest byte lengths and hashes
before serving. A hash establishes integrity relative to the trusted manifest,
not the truth of the source or a digital signature from its publisher.

## Build and check

```powershell
npm run data:build
npm run check
npm run build
```

`data/build_releases.py` produces the public dataset catalogue, immutable
version directories and a small current pointer under `data/generated/datasets`.
It stages new releases before publishing their directory and replaces the
current pointer atomically. Rebuilding identical data preserves its version;
tampering with an existing version fails. The web content build copies only
the public releases and their attribution notice. The final website build
includes `datasets.js` beside the other worker modules.

Each JSONL record includes `releaseSource`, so its release attribution travels
with the record as well as the manifest. Literary records additionally retain
their individual work provenance. The public notice is available beside the
release catalogue. Store it with redistributed association data.

To prepare a private release, supply observation JSONL and review-pack JSON:

```powershell
python data/build_releases.py --observations data/local/reviewed/observations.jsonl --packs data/local/reviewed/packs.json --output data/local/releases
```

Both inputs must be real reviewed files. The command fails for candidate
packs, missing source/annotation redistribution permission, unreviewed source
records, invalid spans, stale text hashes, known unavailable source revisions,
or leaking evaluation groups. It exports a field allowlist: private reviewer
notes and arbitrary quoted context are excluded. See [OBSERVATIONS.md](OBSERVATIONS.md)
for the review contract. The current 30-record study intentionally cannot pass
this commercial-release gate and is not presented as a paid dataset.

## Host configuration

Public datasets run from the existing Cloudflare Pages ASSETS binding when
the built `dist/` is deployed. Private datasets additionally require:

1. `REPERTOIRE` R2 binding. Upload a verified private release directory to
   `datasets/observed-language/<version>/`, then upload its reviewed manifest
   to `datasets/observed-language/current.json` **last**. Never point current
   at partially uploaded files. Test the downloaded bytes against the manifest.
2. `DATA_KEYS` KV binding. Use a cryptographically random bearer key (at least
   32 URL-safe characters). The KV key is the lowercase SHA-256 hex digest of
   the bearer secret. Store JSON with `active: true`, `expiresAt` as a Unix
   timestamp in milliseconds, and `datasets: ["observed-language"]`. A
   `dictionary` scope is separately required for legacy `/api/words/{word}`.
   Keys without explicit scopes no longer access private data.
3. Serve secrets only through approved private channels. Store neither bearer
   secrets nor customer records in static assets, source control or API logs.
   Revoke a key by changing active to false; account for KV propagation in
   operational access-revocation commitments.

To withdraw a private dataset, set `revoked: true` in its current manifest or
remove the pointer. Each API request checks the current manifest and key;
private responses are not cached. This blocks further service of that release.
It does not erase copies already downloaded by clients or discover source
deletions. A rights/source removal must also purge affected raw captures,
revisions, quoted copies, annotations, archived releases, exports and backups;
produce a new cleared version before repointing current. Establish that
retention process before operating an ongoing X collection.

No live Cloudflare deployment, customer key or paid source subscription was
configured during this work. Tests exercise the same worker locally with
real public release files and synthetic private fixtures. Payment activation
and promised availability need their own operational evidence.
