Skip to content

How to Use Laya: Classify Text in a pandas DataFrame with the Open-Source Jev Alternative

Updated on

Install Laya with pip, label a DataFrame text column with typed choice, score, and yes/no questions, and batch it at 13 ms per row. Tested on 40 English, Chinese, Japanese, and Korean tickets, with measured accuracy, speed, and the traps the README doesn't warn about.

The short version: install it, describe the labels you want, and let Laya fill new DataFrame columns:

python -m pip install laya
import pandas as pd
from laya import Router
 
df = pd.DataFrame({"text": [
    "I was charged twice for my September invoice. Please refund the extra payment.",
    "The dashboard has been down since 9am, nobody can log in.",
]})
 
questions = {
    "department": {"type": "choice", "instructions": "Which team should handle this support ticket?",
                   "criteria": {"billing": "charges, invoices, refunds",
                                "technical": "bugs, errors, outages",
                                "account": "login, password, two-factor authentication",
                                "other": "everything else"}},
    "refund_requested": {"type": "noul", "instructions": "Does the customer ask for money back?"},
}
 
router = Router()
results = router.predict_batch([{"state": t, "questions": questions} for t in df["text"]])
df["department"] = [r["answers"]["department"]["choice"] for r in results]
df["refund_p"] = [r["answers"]["refund_requested"]["noul"] for r in results]
print(df)

The first run downloads an ~800 MB checkpoint. After that, each row takes milliseconds, and nothing is generated, so there is no output to parse.

What I found when I tested it on 40 hand-labeled support tickets (10 each in English, Chinese, Japanese, and Korean) on an Apple M4 Max, Laya 0.3.20, September 27, 2026:

QuestionResult
Department (4-way choice)75% overall: English 90%, Korean 80%, Japanese 70%, Chinese 60%
Refund requested (yes/no noul)100% in every language
Urgency (3-level score)30%. The multilingual checkpoint answered "most urgent" for every non-English ticket. Asking a yes/no question instead raised this to 85%
Speed, batched (GPU / Apple MPS)13 ms per ticket with 3 questions
Speed, batched (CPU only)147 ms per ticket

The verdict: Laya is a fast, free first-pass labeler for clear-cut categories and yes/no flags. It is not yet a trustworthy judge of nuance in CJK languages, and its confidence scores do not tell you when it is wrong. The rest of this guide shows the full workflow, the numbers, and how to work around each trap.

What Laya is and why people are talking about it

Laya is an open-source (Apache-2.0) Python library that answers typed questions about a piece of text in a single forward pass. You give it a "state" (a ticket, email, review, or JSON document) and questions of three kinds:

  • choice: pick one label from a set you define ("billing / technical / account / other")
  • score: place the text on an ordered scale ("not urgent / soon / blocking")
  • noul: a yes/no probability ("does the customer threaten to cancel?")

Under the hood it is an encoder model (ModernBERT or mmBERT), not a chat model. It never generates text, so it cannot ramble, hallucinate a label that doesn't exist, or return malformed JSON. The trade-off is that it only does classification.

The timing explains the attention. On September 15, 2026, TypeSafe announced Jev, a hosted "System One" decision model built on the same idea. The launch reached 1,984 points on Hacker News. Four days later, Laya's author posted it as an open alternative they had built a year earlier (1,358 points on HN). The GitHub repo went from about 3k stars to over 26k in a week. Ports and wrappers followed: Kev, OpenJev, and Ollaya, an "Ollama for Jev-style models". Japanese developers on Zenn and Korean developers on GeekNews have been writing about it daily.

For a data person, the practical question is narrower: can this replace an LLM call, or a hand-built classifier, for labeling a text column? That is what the tests below answer.

Step 1 — Install

Laya needs Python 3.10 or newer. It pulls in PyTorch 2.14 and Transformers 5.x. Use a fresh virtual environment so it does not fight your existing torch version:

python3 -m venv .venv
.venv/bin/python -m pip install laya
.venv/bin/python -I -c "import laya; print(laya.__version__)"

For Windows, CUDA-specific torch builds, and Intel GPUs, see the installation section of the README (opens in a new tab). If virtual environments are new to you, start with our Python virtual environment guide.

Plan for disk and download time. Model weights download on first use:

CheckpointEncoderUsed forSize on disk (measured)
layaModernBERT-large, 421M paramsEnglish~807 MB
laya-multilingualmmBERT-base, 322M params100+ languages~1.4 GB for these two combined
laya-typed-decisionsModernBERT-large, 421M paramsbusiness-workflow decisions(included above)
All three together2.2 GB

Without a Hugging Face token, the first checkpoint took 5 minutes 35 seconds to download on my connection. Set a token (export HF_TOKEN=...) to avoid the anonymous rate limit. Set HF_HOME if your home directory is short on space.

Optional extras: laya[serve] (local web playground and HTTP API), laya[mcp] (MCP server), laya[langchain], laya[onnx].

Step 2 — Run your first prediction

The Router is the recommended entry point. It detects the language of each input and sends English text to the English checkpoint and everything else to the multilingual one:

from laya import Router
 
router = Router()
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
 
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"])   # billing
print(result["answers"]["churn_risk"]["noul"])     # 0.879
print(result["routing"]["model"])                  # english

Each answer is a dict. These are the fields you will use most:

FieldMeaning
choice (choice questions)The winning label
probabilitiesProbability for every label or level
noul (yes/no questions)Probability of "yes"
answer_confidenceProbability of the returned answer
result["routing"]Which checkpoint ran and why ("English Latin text", "non-Latin script…")

The warm second call took 33 ms on my machine.

You may see RuntimeWarning: laya: this checkpoint ships invalid temperatures… the first time a checkpoint loads. The library is telling you some confidence values are uncalibrated. Predictions still run, but read the confidence section below before you rely on the numbers.

Step 3 — Label a whole DataFrame column

Do not call predict() in a loop over rows. predict_batch() routes every row first, groups rows by checkpoint and question set, and shares forward passes:

import pandas as pd
from laya import Router
 
df = pd.read_csv("tickets.csv")          # any DataFrame with a text column
 
questions = {
    "department": {"type": "choice", "instructions": "Which team should handle this support ticket?",
                   "criteria": {"billing": "charges, invoices, receipts, refunds, cancelling a paid plan",
                                "technical": "bugs, errors, outages, features not working",
                                "account": "login, password, two-factor authentication, profile settings",
                                "other": "sales questions, discounts, feedback, everything else"}},
    "blocking": {"type": "noul",
                 "instructions": "Is the customer blocked from working right now, or asking for an immediate fix?"},
    "refund_requested": {"type": "noul", "instructions": "Does the customer ask for money back?"},
}
 
router = Router(preload=True)            # load both routed checkpoints up front
results = router.predict_batch(
    [{"state": text, "questions": questions} for text in df["text"]],
    batch_size=32,
)
 
answers = [r["answers"] for r in results]
df["department"] = [a["department"]["choice"] for a in answers]
df["department_conf"] = [a["department"]["answer_confidence"] for a in answers]
df["blocking_p"] = [a["blocking"]["noul"] for a in answers]
df["refund_p"] = [a["refund_requested"]["noul"] for a in answers]
df["checkpoint"] = [r["routing"]["model"] for r in results]

Results come back in input order, so the list assigns straight into columns. If your rows are dicts (email with subject and body, for example), pass the dict as state. Laya accepts JSON-like input.

Measured speed on 40 tickets, 3 questions each, Apple M4 Max:

MethodPer ticket
predict_batch(), auto routing, warm13 ms
predict_batch(), multilingual checkpoint only10 ms
predict_batch(), English checkpoint only34 ms
predict() in a Python loop31 ms
predict_batch() on CPU (Router(device="cpu"))147 ms
First batch that triggers a checkpoint downloadminutes

At 13 ms per row, 100,000 rows take about 22 minutes on a laptop GPU. An LLM API labeling the same rows would cost real money and take longer. On CPU, expect roughly ten times slower. That is still fine for a few thousand rows.

Step 4 — Check accuracy on your own data (what I measured)

Benchmarks in the README use large public datasets. I wanted to know how Laya handles the messy, short texts a support or product team actually has, in the languages this site's readers work in. So I wrote 40 tickets and labeled them by hand: 10 per language, with the same ten situations in each (double charge, outage, email change, student discount, export error, cancellation with refund, password reset, invoice request, praise, 2FA lockout). 40 rows is a smoke test, not a benchmark. Run the same check on 50 to 100 rows of your own data before trusting any number here.

Department (4-way choice) accuracy by checkpoint:

LanguageRouter (auto)English checkpoint forcedMultilingual forcedTyped-decisions forced
English90%90%90%90%
Korean80%30%80%40%
Japanese70%70%70%60%
Chinese60%80%60%70%
Overall75%68%75%65%

Three things stand out:

  1. Let the Router route. Forcing the English checkpoint on Korean dropped accuracy to 30%, and on Chinese it happened to do better. There is no single checkpoint that wins everywhere, and auto routing was the best overall.
  2. Clear-cut categories are easy; "other" is hard. Every language got the double-charge, outage, and export-error tickets right. The misses clustered on "Do you offer a student discount?" (predicted billing in zh / ja / ko) and "the new design looks great" (predicted technical). Both are reasonable-sounding mistakes for a model keyed on product words.
  3. 2FA lockouts went to technical in Chinese, Japanese, and Korean, even though the account criteria mention two-factor authentication explicitly.

The yes/no question was the strongest. "Does the customer ask for money back?" scored 100% in all four languages.

Chinese results match an independent report. A user who tested 20 Chinese routing requests in issue #124 (opens in a new tab) got 70% on the multilingual checkpoint.

Laya vs Jev. Issue #450 (opens in a new tab) ran the English checkpoint on 741 real, outcome-graded government procurement notices. It measured Laya at 0.780 accuracy against Jev's 0.919.

The traps, and how to work around them

Trap 1: score questions on the multilingual checkpoint collapse to the last option

I asked for urgency on a three-level scale (not urgent, soon, blocking work right now). The multilingual checkpoint chose blocking work right now for all 40 tickets, including "the new charts look great, thanks". With auto routing, only the 10 English tickets escaped, and overall urgency accuracy was 30%.

This is a known, open bug. Issue #131 (opens in a new tab) reports that laya-multilingual never picks the first-listed score level (0 of 290 in English, 0 of 300 in Japanese), and the probability mass follows the slot, not the label.

Workaround that worked: turn the scale into yes/no questions. Replacing the urgency score with one noul, "Is the customer blocked from working right now, or asking for an immediate fix?", scored 85% (English 90%, Chinese 90%, Korean 90%, Japanese 70%). If you need three levels, ask two yes/no questions ("blocking now?" and "needs a reply today?") and combine them in pandas.

Trap 2: Confidence does not separate right from wrong

The obvious plan is to auto-accept confident rows and send the rest to a human. On my run the gate helped only a little:

Accept rows with answer_confidence ≥Rows auto-acceptedAccuracy on accepted rowsWrong rows let through
0.588%74%9
0.768%78%6
0.955%82%4

Two of the worst errors came with 0.99 confidence: the praise ticket labeled technical in Chinese, and the discount question labeled billing in Japanese. The load-time temperature warning and issue #124 ("confident wrong answers persist") describe the same thing.

What to do: treat confidence as a weak signal. Spot-check a random sample per label, not just the low-confidence tail. If a category matters (legal complaints, churn threats), give it its own noul question. The yes/no questions were far more reliable than the 4-way choice.

Trap 3: Translating the questions did not reliably help

I rewrote the question and every label description in Chinese, Japanese, and Korean (keeping the English label keys) and reran the non-English tickets. Chinese went from 60% to 80%, while Japanese dropped from 70% to 60% and Korean from 80% to 70%. The overall result was unchanged at 70%. Keep the questions in English unless your own test says otherwise.

Trap 4: Short Latin-script text is routed to the English checkpoint

The Router needs enough text to detect a language. The README notes that very short Spanish or Portuguese strings ("Esqueci minha senha") go to the default checkpoint, which is English. If most of your data is not English, change the default:

router = Router(default="multilingual")

You can inspect a routing decision without running the model:

router.route({"body": "Der Kunde wurde zweimal belastet"}).reason

Trap 5: Long documents are cut off silently

The multilingual checkpoint ships with a 1,024-token limit, and the English one reads 512 tokens. For long emails or documents, pass max_len and name the checkpoint. According to the README, accuracy holds up to about 4,000 tokens and varies beyond that:

result = router.predict(long_document, questions, model="multilingual", max_len=8192)

Trap 6: Memory churn when languages alternate

Router() keeps two checkpoints resident by default. With max_loaded=1, every switch between English and non-English rows reloads a model, which the README measured at 7 to 10 seconds per switch. For mixed-language DataFrames, use Router(preload=True) or keep the default, and always use predict_batch() so rows are grouped by checkpoint.

Laya, Jev, or an LLM?

LayaJev (TypeSafe)General LLM (API)
Runs locally, freeYes, Apache-2.0No, hosted APIOnly with a local model
Speed per row~10–35 ms on GPUHostedHundreds of ms to seconds
Output formatAlways a valid label or probabilityTyped answersNeeds parsing or structured output
Accuracy on nuanced / CJK textMixed (75% in my 4-language test)Higher in the #450 comparison (0.919 vs 0.780, English)Usually best
Setuppip install, ~800 MB per checkpointAccount and API keyAPI key

A sensible pattern: run Laya over the whole column first. Send only the rows that fall into weak categories (here, other and anything involving accounts in CJK text) or that fail your spot checks to an LLM or a human. You pay LLM prices on a fraction of the data.

If the labels need to be domain decisions rather than generic ones, the README's fine-tuning notebook is the next step. The authors report 0.766 vs 0.362 accuracy after fine-tuning on their typed-decisions benchmark. I did not test fine-tuning.

Step 5 — Explore the labeled data visually

Once the new columns exist, the useful questions are about distribution: which departments get the most tickets, which languages produce low-confidence labels, and where refund requests cluster. PyGWalker (opens in a new tab) turns the DataFrame into a drag-and-drop chart builder inside Jupyter:

import pygwalker as pyg
 
pyg.walk(df[["text", "department", "department_conf", "blocking_p", "refund_p", "checkpoint"]])

Put department on the x-axis and count on the y-axis, color by checkpoint, and the routing pattern shows up in one chart. Filter department_conf < 0.7 to build the human-review queue. Outside Jupyter, pyg.to_html(df) writes a standalone HTML page (tested with pygwalker 0.5.0.1). For a deeper walkthrough, see the PyGWalker quick start.

If you do this kind of labeling inside notebooks regularly, RunCell (opens in a new tab) is an AI agent that works in JupyterLab against the live DataFrame. Asking it to "add Laya labels to df and chart the low-confidence rows" is a natural fit, and it can see the actual column values instead of guessing.

What I tested and what comes from docs

  • Tested (Sept 27, 2026, macOS on Apple M4 Max, Python 3.12, laya 0.3.20, torch 2.14.0, transformers 5.17.0):
    • install
    • first prediction and output format
    • checkpoint download time and disk size
    • predict() vs predict_batch() speed on MPS and CPU
    • the 40-ticket accuracy tables for all three checkpoints plus auto routing
    • the score-collapse bug and the noul workaround
    • confidence gating
    • translated questions
    • PyGWalker to_html on the results
  • From the README and GitHub issues, not tested here:
    • long-document accuracy with max_len=8192
    • T4 GPU latency figures
    • default="multilingual" routing behavior
    • memory-churn timings
    • fine-tuning results
    • the Jev comparison in issue #450

The project ships almost daily. Check the release notes (opens in a new tab) and issue #131 before assuming the score behavior above still applies.

FAQ

Related Guides