Skip to content

Laya vs Jev, With the Source Limits Attached

Jev is TypeSafe’s hosted System-1 API. Laya is Convai’s open-weight model family under Apache 2.0. This page does not declare a winner. It repeats figures the Laya maintainers published, and it marks which of those figures they did not measure themselves.

Source: the “Laya (with routing) vs TypeSafe Jev” section of the Hugging Face model card and the GitHub README, checked 29 Sep 2026. This site did not re-run the benchmarks. Jev cells are third-party published numbers. The README says sample sizes and prompts differ, so they are not a controlled head-to-head. Jev latency citations in that README point at AbdelStark/jev-benchmarks and nibzard/decision-model-benchmark.

Published comparison

FigureJev 1.13.0Laya as citedWhat the number is
typed-decisions, 2,000 decisions0.727 published0.766Fine-tuned laya-typed-decisions, not the base checkpoint you get from a plain install
AG News, 4 labels0.9100.950Upstream routed comparison
DAIR Emotion, 6 labels0.4800.595Upstream routed comparison. The README also says Jev assigned zero probability to the true label on 16% of examples
Banking770.870 on 72 labels0.425 on 77 labelsJev leads. The label counts differ, and Laya’s figure is at the default head budget
ECE, lower is better0.246 in the headline table0.081Laya’s 0.081 is after temperature fitting. Raw ECE before that fit favors Jev on the typed-decisions table (0.144 vs 0.213)
p50 latency, 1 question236–276 ms, third-party32.8 msLaya figure is a T4 measurement in the upstream write-up, not a latency promise for laya-model.online
WeightsClosed APIApache 2.0Open weights can be run on your own machine. This website’s live button still uses the hosted Space
Listed token price$0.042 / 1M tokens in the Laya README$0 self-hostedSelf-hosting is not the same as this free single-message website

Where the base Laya checkpoints are weak

On the same 2,000 typed decisions, the English base is 0.362 and the multilingual base is reported as 0.342 in the checkpoint table and 0.352 in the later honest-limits / PyPI 0.3.21 wording. The per-question majority baseline is 0.461. The upstream card’s own summary: Laya is a fast base to specialise, not a zero-shot decision engine. Quoting 0.766 without that sentence overstates what an unfine-tuned checkpoint does.

Soft accuracy on typed-decisions is another Jev lead in the README: 0.580 versus 0.471 for the fine-tuned Laya checkpoint, against the teacher’s full probability distribution. Hard-label accuracy and soft match are different questions.

What this means if you are routing tickets

Neither column is a measured accuracy for the templates on this site. Our ticket and feedback pages publish six individual Space calls, including a negation the model missed. Use the ticket classifier to see a live label, then evaluate it on your own sample before you connect it to a queue.

The wire protocol is shared. Laya’s laya-serve speaks POST /v1/systemone, the same shape Jev uses. This website does not expose that server. It calls the public Space playground.