Laya vs Jev, With the Source Limits Attached
Jev is TypeSafe’s hosted System-1 API. Laya is Convai’s open-weight model family under Apache 2.0. This page does not declare a winner. It repeats figures the Laya maintainers published, and it marks which of those figures they did not measure themselves.
Source: the “Laya (with routing) vs TypeSafe Jev” section of the Hugging Face model card and the GitHub README, checked 29 Sep 2026. This site did not re-run the benchmarks. Jev cells are third-party published numbers. The README says sample sizes and prompts differ, so they are not a controlled head-to-head. Jev latency citations in that README point at AbdelStark/jev-benchmarks and nibzard/decision-model-benchmark.
Published comparison
| Figure | Jev 1.13.0 | Laya as cited | What the number is |
|---|---|---|---|
| typed-decisions, 2,000 decisions | 0.727 published | 0.766 | Fine-tuned laya-typed-decisions, not the base checkpoint you get from a plain install |
| AG News, 4 labels | 0.910 | 0.950 | Upstream routed comparison |
| DAIR Emotion, 6 labels | 0.480 | 0.595 | Upstream routed comparison. The README also says Jev assigned zero probability to the true label on 16% of examples |
| Banking77 | 0.870 on 72 labels | 0.425 on 77 labels | Jev leads. The label counts differ, and Laya’s figure is at the default head budget |
| ECE, lower is better | 0.246 in the headline table | 0.081 | Laya’s 0.081 is after temperature fitting. Raw ECE before that fit favors Jev on the typed-decisions table (0.144 vs 0.213) |
| p50 latency, 1 question | 236–276 ms, third-party | 32.8 ms | Laya figure is a T4 measurement in the upstream write-up, not a latency promise for laya-model.online |
| Weights | Closed API | Apache 2.0 | Open weights can be run on your own machine. This website’s live button still uses the hosted Space |
| Listed token price | $0.042 / 1M tokens in the Laya README | $0 self-hosted | Self-hosting is not the same as this free single-message website |
Where the base Laya checkpoints are weak
On the same 2,000 typed decisions, the English base is 0.362 and the multilingual base is reported as 0.342 in the checkpoint table and 0.352 in the later honest-limits / PyPI 0.3.21 wording. The per-question majority baseline is 0.461. The upstream card’s own summary: Laya is a fast base to specialise, not a zero-shot decision engine. Quoting 0.766 without that sentence overstates what an unfine-tuned checkpoint does.
Soft accuracy on typed-decisions is another Jev lead in the README: 0.580 versus 0.471 for the fine-tuned Laya checkpoint, against the teacher’s full probability distribution. Hard-label accuracy and soft match are different questions.
What this means if you are routing tickets
Neither column is a measured accuracy for the templates on this site. Our ticket and feedback pages publish six individual Space calls, including a negation the model missed. Use the ticket classifier to see a live label, then evaluate it on your own sample before you connect it to a queue.
The wire protocol is shared. Laya’s laya-serve speaks POST /v1/systemone, the same shape Jev uses. This website does not expose that server. It calls the public Space playground.