How to Use Laya: Classify Text in a pandas DataFrame with the Open-Source Jev Alternative
Updated on

The short version: install it, describe the labels you want, and let Laya fill new DataFrame columns:
python -m pip install layaimport pandas as pd
from laya import Router
df = pd.DataFrame({"text": [
"I was charged twice for my September invoice. Please refund the extra payment.",
"The dashboard has been down since 9am, nobody can log in.",
]})
questions = {
"department": {"type": "choice", "instructions": "Which team should handle this support ticket?",
"criteria": {"billing": "charges, invoices, refunds",
"technical": "bugs, errors, outages",
"account": "login, password, two-factor authentication",
"other": "everything else"}},
"refund_requested": {"type": "noul", "instructions": "Does the customer ask for money back?"},
}
router = Router()
results = router.predict_batch([{"state": t, "questions": questions} for t in df["text"]])
df["department"] = [r["answers"]["department"]["choice"] for r in results]
df["refund_p"] = [r["answers"]["refund_requested"]["noul"] for r in results]
print(df)The first run downloads an ~800 MB checkpoint. After that, each row takes milliseconds, and nothing is generated, so there is no output to parse.
What I found when I tested it on 40 hand-labeled support tickets (10 each in English, Chinese, Japanese, and Korean) on an Apple M4 Max, Laya 0.3.20, September 27, 2026:
| Question | Result |
|---|---|
Department (4-way choice) | 75% overall: English 90%, Korean 80%, Japanese 70%, Chinese 60% |
Refund requested (yes/no noul) | 100% in every language |
Urgency (3-level score) | 30%. The multilingual checkpoint answered "most urgent" for every non-English ticket. Asking a yes/no question instead raised this to 85% |
| Speed, batched (GPU / Apple MPS) | 13 ms per ticket with 3 questions |
| Speed, batched (CPU only) | 147 ms per ticket |
The verdict: Laya is a fast, free first-pass labeler for clear-cut categories and yes/no flags. It is not yet a trustworthy judge of nuance in CJK languages, and its confidence scores do not tell you when it is wrong. The rest of this guide shows the full workflow, the numbers, and how to work around each trap.
What Laya is and why people are talking about it
Laya is an open-source (Apache-2.0) Python library that answers typed questions about a piece of text in a single forward pass. You give it a "state" (a ticket, email, review, or JSON document) and questions of three kinds:
choice: pick one label from a set you define ("billing / technical / account / other")score: place the text on an ordered scale ("not urgent / soon / blocking")noul: a yes/no probability ("does the customer threaten to cancel?")
Under the hood it is an encoder model (ModernBERT or mmBERT), not a chat model. It never generates text, so it cannot ramble, hallucinate a label that doesn't exist, or return malformed JSON. The trade-off is that it only does classification.
The timing explains the attention. On September 15, 2026, TypeSafe announced Jev, a hosted "System One" decision model built on the same idea. The launch reached 1,984 points on Hacker News. Four days later, Laya's author posted it as an open alternative they had built a year earlier (1,358 points on HN). The GitHub repo went from about 3k stars to over 26k in a week. Ports and wrappers followed: Kev, OpenJev, and Ollaya, an "Ollama for Jev-style models". Japanese developers on Zenn and Korean developers on GeekNews have been writing about it daily.
For a data person, the practical question is narrower: can this replace an LLM call, or a hand-built classifier, for labeling a text column? That is what the tests below answer.
- How to Use Laya: Classify Text in a pandas DataFrame with the Open-Source Jev Alternative
- Claude Code Reads AGENTS.md Now: Setup, Loading Modes, and Sharing One File with Codex, OpenCode, and Cursor
- How to View Deleted Reddit Posts (2026 Guide): 5 Ways That Still Work
- GPT Image 2.5: How to Use It, Flare vs Sunburst, and API Pricing
- How to Use DeepSeek Harness: Install, Set Up, and Run Your First Agent
- Runcell Science: An Open Source Alternative to Claude Science for Research Workflows
- How to Make Mac Not Sleep: Keep Codex, Claude Code, and AI Agents Running
- OpenClaw vs ZeroClaw vs Pi Agent vs Nanobot: Which AI Agent Stack Should You Choose in 2026?
- Can Claude Code Analyze Jupyter Notebooks for Data Science? What It Actually Does
- Claude Code Routines: Why AI Agent Cron Jobs Matter
- Claude Code Desktop Bypass Permissions: How to Enable It
- How to Build Two Python Agents with Google’s A2A Protocol - Step by Step Tutorial
- Top 10 growing data visualization libraries in Python in 2025
Step 1 — Install
Laya needs Python 3.10 or newer. It pulls in PyTorch 2.14 and Transformers 5.x. Use a fresh virtual environment so it does not fight your existing torch version:
python3 -m venv .venv
.venv/bin/python -m pip install laya
.venv/bin/python -I -c "import laya; print(laya.__version__)"For Windows, CUDA-specific torch builds, and Intel GPUs, see the installation section of the README (opens in a new tab). If virtual environments are new to you, start with our Python virtual environment guide.
Plan for disk and download time. Model weights download on first use:
| Checkpoint | Encoder | Used for | Size on disk (measured) |
|---|---|---|---|
laya | ModernBERT-large, 421M params | English | ~807 MB |
laya-multilingual | mmBERT-base, 322M params | 100+ languages | ~1.4 GB for these two combined |
laya-typed-decisions | ModernBERT-large, 421M params | business-workflow decisions | (included above) |
| All three together | 2.2 GB |
Without a Hugging Face token, the first checkpoint took 5 minutes 35 seconds to download on my connection. Set a token (export HF_TOKEN=...) to avoid the anonymous rate limit. Set HF_HOME if your home directory is short on space.
Optional extras: laya[serve] (local web playground and HTTP API), laya[mcp] (MCP server), laya[langchain], laya[onnx].
Step 2 — Run your first prediction
The Router is the recommended entry point. It detects the language of each input and sends English text to the English checkpoint and everything else to the multilingual one:
from laya import Router
router = Router()
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"]) # billing
print(result["answers"]["churn_risk"]["noul"]) # 0.879
print(result["routing"]["model"]) # englishEach answer is a dict. These are the fields you will use most:
| Field | Meaning |
|---|---|
choice (choice questions) | The winning label |
probabilities | Probability for every label or level |
noul (yes/no questions) | Probability of "yes" |
answer_confidence | Probability of the returned answer |
result["routing"] | Which checkpoint ran and why ("English Latin text", "non-Latin script…") |
The warm second call took 33 ms on my machine.
You may see RuntimeWarning: laya: this checkpoint ships invalid temperatures… the first time a checkpoint loads. The library is telling you some confidence values are uncalibrated. Predictions still run, but read the confidence section below before you rely on the numbers.
Step 3 — Label a whole DataFrame column
Do not call predict() in a loop over rows. predict_batch() routes every row first, groups rows by checkpoint and question set, and shares forward passes:
import pandas as pd
from laya import Router
df = pd.read_csv("tickets.csv") # any DataFrame with a text column
questions = {
"department": {"type": "choice", "instructions": "Which team should handle this support ticket?",
"criteria": {"billing": "charges, invoices, receipts, refunds, cancelling a paid plan",
"technical": "bugs, errors, outages, features not working",
"account": "login, password, two-factor authentication, profile settings",
"other": "sales questions, discounts, feedback, everything else"}},
"blocking": {"type": "noul",
"instructions": "Is the customer blocked from working right now, or asking for an immediate fix?"},
"refund_requested": {"type": "noul", "instructions": "Does the customer ask for money back?"},
}
router = Router(preload=True) # load both routed checkpoints up front
results = router.predict_batch(
[{"state": text, "questions": questions} for text in df["text"]],
batch_size=32,
)
answers = [r["answers"] for r in results]
df["department"] = [a["department"]["choice"] for a in answers]
df["department_conf"] = [a["department"]["answer_confidence"] for a in answers]
df["blocking_p"] = [a["blocking"]["noul"] for a in answers]
df["refund_p"] = [a["refund_requested"]["noul"] for a in answers]
df["checkpoint"] = [r["routing"]["model"] for r in results]Results come back in input order, so the list assigns straight into columns. If your rows are dicts (email with subject and body, for example), pass the dict as state. Laya accepts JSON-like input.
Measured speed on 40 tickets, 3 questions each, Apple M4 Max:
| Method | Per ticket |
|---|---|
predict_batch(), auto routing, warm | 13 ms |
predict_batch(), multilingual checkpoint only | 10 ms |
predict_batch(), English checkpoint only | 34 ms |
predict() in a Python loop | 31 ms |
predict_batch() on CPU (Router(device="cpu")) | 147 ms |
| First batch that triggers a checkpoint download | minutes |
At 13 ms per row, 100,000 rows take about 22 minutes on a laptop GPU. An LLM API labeling the same rows would cost real money and take longer. On CPU, expect roughly ten times slower. That is still fine for a few thousand rows.
Step 4 — Check accuracy on your own data (what I measured)
Benchmarks in the README use large public datasets. I wanted to know how Laya handles the messy, short texts a support or product team actually has, in the languages this site's readers work in. So I wrote 40 tickets and labeled them by hand: 10 per language, with the same ten situations in each (double charge, outage, email change, student discount, export error, cancellation with refund, password reset, invoice request, praise, 2FA lockout). 40 rows is a smoke test, not a benchmark. Run the same check on 50 to 100 rows of your own data before trusting any number here.
Department (4-way choice) accuracy by checkpoint:
| Language | Router (auto) | English checkpoint forced | Multilingual forced | Typed-decisions forced |
|---|---|---|---|---|
| English | 90% | 90% | 90% | 90% |
| Korean | 80% | 30% | 80% | 40% |
| Japanese | 70% | 70% | 70% | 60% |
| Chinese | 60% | 80% | 60% | 70% |
| Overall | 75% | 68% | 75% | 65% |
Three things stand out:
- Let the Router route. Forcing the English checkpoint on Korean dropped accuracy to 30%, and on Chinese it happened to do better. There is no single checkpoint that wins everywhere, and auto routing was the best overall.
- Clear-cut categories are easy; "other" is hard. Every language got the double-charge, outage, and export-error tickets right. The misses clustered on "Do you offer a student discount?" (predicted
billingin zh / ja / ko) and "the new design looks great" (predictedtechnical). Both are reasonable-sounding mistakes for a model keyed on product words. - 2FA lockouts went to
technicalin Chinese, Japanese, and Korean, even though theaccountcriteria mention two-factor authentication explicitly.
The yes/no question was the strongest. "Does the customer ask for money back?" scored 100% in all four languages.
Chinese results match an independent report. A user who tested 20 Chinese routing requests in issue #124 (opens in a new tab) got 70% on the multilingual checkpoint.
Laya vs Jev. Issue #450 (opens in a new tab) ran the English checkpoint on 741 real, outcome-graded government procurement notices. It measured Laya at 0.780 accuracy against Jev's 0.919.
The traps, and how to work around them
Trap 1: score questions on the multilingual checkpoint collapse to the last option
I asked for urgency on a three-level scale (not urgent, soon, blocking work right now). The multilingual checkpoint chose blocking work right now for all 40 tickets, including "the new charts look great, thanks". With auto routing, only the 10 English tickets escaped, and overall urgency accuracy was 30%.
This is a known, open bug. Issue #131 (opens in a new tab) reports that laya-multilingual never picks the first-listed score level (0 of 290 in English, 0 of 300 in Japanese), and the probability mass follows the slot, not the label.
Workaround that worked: turn the scale into yes/no questions. Replacing the urgency score with one noul, "Is the customer blocked from working right now, or asking for an immediate fix?", scored 85% (English 90%, Chinese 90%, Korean 90%, Japanese 70%). If you need three levels, ask two yes/no questions ("blocking now?" and "needs a reply today?") and combine them in pandas.
Trap 2: Confidence does not separate right from wrong
The obvious plan is to auto-accept confident rows and send the rest to a human. On my run the gate helped only a little:
Accept rows with answer_confidence ≥ | Rows auto-accepted | Accuracy on accepted rows | Wrong rows let through |
|---|---|---|---|
| 0.5 | 88% | 74% | 9 |
| 0.7 | 68% | 78% | 6 |
| 0.9 | 55% | 82% | 4 |
Two of the worst errors came with 0.99 confidence: the praise ticket labeled technical in Chinese, and the discount question labeled billing in Japanese. The load-time temperature warning and issue #124 ("confident wrong answers persist") describe the same thing.
What to do: treat confidence as a weak signal. Spot-check a random sample per label, not just the low-confidence tail. If a category matters (legal complaints, churn threats), give it its own noul question. The yes/no questions were far more reliable than the 4-way choice.
Trap 3: Translating the questions did not reliably help
I rewrote the question and every label description in Chinese, Japanese, and Korean (keeping the English label keys) and reran the non-English tickets. Chinese went from 60% to 80%, while Japanese dropped from 70% to 60% and Korean from 80% to 70%. The overall result was unchanged at 70%. Keep the questions in English unless your own test says otherwise.
Trap 4: Short Latin-script text is routed to the English checkpoint
The Router needs enough text to detect a language. The README notes that very short Spanish or Portuguese strings ("Esqueci minha senha") go to the default checkpoint, which is English. If most of your data is not English, change the default:
router = Router(default="multilingual")You can inspect a routing decision without running the model:
router.route({"body": "Der Kunde wurde zweimal belastet"}).reasonTrap 5: Long documents are cut off silently
The multilingual checkpoint ships with a 1,024-token limit, and the English one reads 512 tokens. For long emails or documents, pass max_len and name the checkpoint. According to the README, accuracy holds up to about 4,000 tokens and varies beyond that:
result = router.predict(long_document, questions, model="multilingual", max_len=8192)Trap 6: Memory churn when languages alternate
Router() keeps two checkpoints resident by default. With max_loaded=1, every switch between English and non-English rows reloads a model, which the README measured at 7 to 10 seconds per switch. For mixed-language DataFrames, use Router(preload=True) or keep the default, and always use predict_batch() so rows are grouped by checkpoint.
Laya, Jev, or an LLM?
| Laya | Jev (TypeSafe) | General LLM (API) | |
|---|---|---|---|
| Runs locally, free | Yes, Apache-2.0 | No, hosted API | Only with a local model |
| Speed per row | ~10–35 ms on GPU | Hosted | Hundreds of ms to seconds |
| Output format | Always a valid label or probability | Typed answers | Needs parsing or structured output |
| Accuracy on nuanced / CJK text | Mixed (75% in my 4-language test) | Higher in the #450 comparison (0.919 vs 0.780, English) | Usually best |
| Setup | pip install, ~800 MB per checkpoint | Account and API key | API key |
A sensible pattern: run Laya over the whole column first. Send only the rows that fall into weak categories (here, other and anything involving accounts in CJK text) or that fail your spot checks to an LLM or a human. You pay LLM prices on a fraction of the data.
If the labels need to be domain decisions rather than generic ones, the README's fine-tuning notebook is the next step. The authors report 0.766 vs 0.362 accuracy after fine-tuning on their typed-decisions benchmark. I did not test fine-tuning.
Step 5 — Explore the labeled data visually
Once the new columns exist, the useful questions are about distribution: which departments get the most tickets, which languages produce low-confidence labels, and where refund requests cluster. PyGWalker (opens in a new tab) turns the DataFrame into a drag-and-drop chart builder inside Jupyter:
import pygwalker as pyg
pyg.walk(df[["text", "department", "department_conf", "blocking_p", "refund_p", "checkpoint"]])Put department on the x-axis and count on the y-axis, color by checkpoint, and the routing pattern shows up in one chart. Filter department_conf < 0.7 to build the human-review queue. Outside Jupyter, pyg.to_html(df) writes a standalone HTML page (tested with pygwalker 0.5.0.1). For a deeper walkthrough, see the PyGWalker quick start.
If you do this kind of labeling inside notebooks regularly, RunCell (opens in a new tab) is an AI agent that works in JupyterLab against the live DataFrame. Asking it to "add Laya labels to df and chart the low-confidence rows" is a natural fit, and it can see the actual column values instead of guessing.
What I tested and what comes from docs
- Tested (Sept 27, 2026, macOS on Apple M4 Max, Python 3.12, laya 0.3.20, torch 2.14.0, transformers 5.17.0):
- install
- first prediction and output format
- checkpoint download time and disk size
predict()vspredict_batch()speed on MPS and CPU- the 40-ticket accuracy tables for all three checkpoints plus auto routing
- the score-collapse bug and the
noulworkaround - confidence gating
- translated questions
- PyGWalker
to_htmlon the results
- From the README and GitHub issues, not tested here:
- long-document accuracy with
max_len=8192 - T4 GPU latency figures
default="multilingual"routing behavior- memory-churn timings
- fine-tuning results
- the Jev comparison in issue #450
- long-document accuracy with
The project ships almost daily. Check the release notes (opens in a new tab) and issue #131 before assuming the score behavior above still applies.
FAQ
Related Guides
- pandas apply: Run a Function on Every Row or Column
- pandas String Operations
- Text Cleaning in Python
- Python Virtual Environments
- PyGWalker Quick Start
- Jupyter AI with RunCell