Laya is an open-source decision model released on September 18, 2026 by Convai Innovations, based in India. You give it a “state” such as an email, a support ticket or a JSON object, ask typed questions like “which department should handle this?”, “how urgent is it, on a scale?” or “is the customer asking for a refund?”, and it returns answers with confidence scores in a single forward pass, without generating any text. It is designed for the same job as TypeSafe AI’s decision model “Jev”, with the difference that its weights are public and you can run it in your own environment. It is licensed under Apache 2.0 and installs with a single pip install laya.
Key Features
- Three kinds of typed questions: You can ask in three formats ─ choice (pick one option), score (return an ordinal score) and noul (return Yes/No as a probability). Answers come back as structured values, not prose
- All questions answered in one pass: Instead of generating text token by token, the encoder answers several questions at once in a single forward pass. According to the official README, it takes roughly 33 to 40 ms per question on a T4 GPU, and about 7 ms per question when 10 questions are batched
- Calibrated confidence: Reinforcement learning and temperature scaling tune the reported confidence to track actual accuracy. The aim is that answers given with 0.9 confidence are right about 90% of the time, which suits workflows where you set a threshold and send the rest to human review
- Automatic language routing: It detects the script and language of the input and routes it to the English or multilingual model automatically. You can also plug in your own language-detection function
- Three checkpoints for different uses: Choose among laya for English (ModernBERT-large, 421M parameters), laya-multilingual for 100+ languages (mmBERT-base, 322M parameters), and laya-typed-decisions, further fine-tuned for business decisions
- Jev-compatible HTTP server:
pip install "laya[serve]"gives you an API server that speaks the samePOST /v1/systemoneformat as Jev. Clients already using Jev are meant to migrate just by changing the endpoint URL
Pricing
| Offering | Price | Details |
|---|---|---|
| Self-hosted (core) | Free | Apache 2.0. Download the three checkpoints from Hugging Face and run them in your own environment |
| Hosted API (impossibl.com) | No per-token charge for Laya | An API run by a third party. Requires a funded account and an API key |
Pricing information is as of September 2026. The core is open source, so the only cost is the machine you run it on. The hosted API is provided by a third party, not Convai Innovations, and its terms may change. Check the GitHub repository and the impossibl.com documentation for the latest details.
Pros & Cons
✅ Pros
- Answers always come from the predefined options or numeric range, so department names that don’t exist or stray sentences never slip into the results, and you can pass them straight to downstream processing
- At tens of milliseconds per question, it can be built into real-time routing of incoming emails and tickets one by one
- Because it runs on your own server, you can make decisions on sensitive data such as customer email bodies without sending it to an external API
- Calibrated confidence lets you draw the line numerically, for example “humans only review answers below a given confidence”
⚠️ Cons
- According to the “Honest limits” section of the official README, the base checkpoints used as-is on business decisions score close to chance (0.362 versus a random baseline of 0.318). In practice you should plan on fine-tuning
- With more than 20 options, accuracy drops sharply under the default settings. You need to adjust the settings or narrow down the candidates
- The README states that the multilingual checkpoint ships without tuned calibration (temperature), so you have to calibrate it on your own data first. A tendency for multilingual score questions to be swayed by option order is also reported
- It is very new, with updates from v0.1.0 to v0.3.11 within days. The API may change in a short time, and there are few production case studies so far
Comparison with Similar Services
| Service | How it differs from Laya |
|---|---|
| Jev (TypeSafe AI) | The same kind of typed decision API, but a closed, paid service (input tokens cost $0.042 per million via impossibl.com). Laya publishes its weights and runs free in your own environment |
| OpenAI Structured Outputs / Function Calling | Has a generative LLM answer in JSON. Flexible, but slower and billed per token. Laya generates no text and finishes in one pass |
| Hugging Face zero-shot classification (e.g. BART-MNLI) | Mainly single-label classification. Laya answers multiple choice, score and Yes/No questions together, with calibrated confidence |
Who Is It For
- Developers who want to automate department routing and urgency scoring for inquiry emails and support tickets
- Teams that want to run churn-signal detection, refund-request detection or content-toxicity checks at low cost
- People whose large classification workloads are bottlenecked by generative LLM API costs or latency
- Anyone building agent task routing or prompt-injection guardrails with a lightweight model
- Jev users who want to switch to self-hosting while keeping the same call format
Summary
Calling a generative LLM and having it assemble JSON for every decision tends to be heavy in both speed and cost. Laya is a pragmatic alternative that carves out just the decision step and hands it to a small model that generates no text. On the other hand, the project itself states that the base checkpoints are not accurate enough for business decisions out of the box, so it is realistic to evaluate it with fine-tuning and confidence calibration in mind. A good starting point is to try it on a sample of your own emails or tickets and run it with a confidence threshold combined with human review.