# Prediction providers: Jev and Laya

Reviewed 2026-09-22 against primary documentation and upstream source. This is an integration assessment, not a Superb model benchmark.

## Recommendation

Keep Jev as the v1 default for `predict_async()` and offer Laya as an explicit, self-hosted alternative. Do not switch the default on published cross-project accuracy or latency claims. No matched Superb evaluation has been run: this workspace has no configured TypeSafe API key or running Laya model. Local response fixtures establish adapter behavior only, not prediction accuracy, calibration, availability, or model acceptance.

`predict()` must continue to read the accepted Solana consensus record. Neither provider's contextual judgment establishes blockchain consensus or substitutes for a missing accepted record. Provider selection affects `predict_async()` only.

## Interfaces and limits

| Concern | Jev | Laya |
| --- | --- | --- |
| Invocation | Bearer-authenticated `POST https://api.typesafe.ai/v1/systemone`; JSON `model`, `state`, `questions` | Python `Agent.system_one(state, questions)`; `predict` is an alias. Superb needs an HTTP bridge around this Python call. |
| Choice contract | `questions[id]` contains `type: "choice"`, `instructions`, and a map of option keys to descriptions | Same choice input shape; result includes `answers[id].choice`, `probabilities`, and `confidence` |
| Inventory | At most 255 choices, so reserve one for `unclear` and allow at most 254 senses | No equivalent documented fixed safe inventory size: token budgets determine whether descriptions survive |
| Context | Jev 1.13: 64k tokens for the whole request; 32k for state plus longest question | English defaults: 512 total and 192 head tokens. Multilingual/typed-decisions defaults: 1,024 total and 256 head tokens. |

Sources: [TypeSafe API](https://docs.typesafe.ai/api), [Choice](https://docs.typesafe.ai/primitives/choice), [Jev model limits](https://docs.typesafe.ai/models), [Laya runtime](https://raw.githubusercontent.com/NandhaKishorM/laya/main/laya/agent.py), [Laya documented context budgets](https://github.com/NandhaKishorM/laya#honest-limits).

Laya's sequence builder silently limits each rendered option to 48 tokens, can shorten all options further when the head budget is crowded, clips instructions, and retains only the prefix of state that fits. Its runtime detects missing option markers, but that does not detect partially lost descriptions or state. Consequently, receiving probabilities for every sense does not prove the model saw every definition. The bridge should preflight the loaded tokenizer and reject requests that would lose text; do not use character counts as token counts. Keep the passage, target, and span within that checked state. These recommendations follow directly from [sequence construction](https://raw.githubusercontent.com/NandhaKishorM/laya/main/laya/common.py) and the [runtime marker check](https://raw.githubusercontent.com/NandhaKishorM/laya/main/laya/agent.py).

Validate exact option coverage, finite probabilities in range, sum within a rounding tolerance, selected option at the maximum probability, and finite confidence. Laya rounds probabilities to four decimal places; its generic `laya-rl-agent` model string is insufficient to identify a deployed checkpoint. Have the bridge report a pinned checkpoint/revision and relevant configuration in its model identity. [Laya response implementation](https://raw.githubusercontent.com/NandhaKishorM/laya/main/laya/agent.py)

## Evidence required before changing the default

Laya's own documentation says its Jev comparison uses third-party Jev results with different sample sizes and prompts, rather than a matched run. It also describes weaker large-label behavior, overconfidence, and checkpoint-dependent performance. Those results do not establish performance on Superb word senses. [Laya benchmark qualifications](https://github.com/NandhaKishorM/laya#laya-with-routing-vs-jev)

Run both providers on the same held-out, independently labeled Superb sentences, combining internet usages and superb-catalogue examples without source or near-duplicate leakage. Include ambiguous/insufficient contexts, long sentences, long glosses, large inventories, target spans, and each supported language. Measure sense accuracy, abstention coverage and error rate, calibration, truncation rejection, latency including transport, throughput, and actual hosting/API cost. Freeze inventory, prompts, checkpoint/configuration, and split before evaluation; fit any confidence threshold or temperature on a separate validation split.

Choose an acceptance margin and operational budget before running that comparison. Pin the winning version: TypeSafe documents that `jev-latest` can move and recommends version IDs when thresholds depend on a model. [TypeSafe aliases](https://docs.typesafe.ai/models#aliases)
