AI 日报hiw3c.com

Laya开发人员指南:零镜头决策和校准

原文标题 · A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration
MarkTechPost www.marktechpost.com RSS 全文
正文为英文,可一键机器翻译(仅首次需要等待)

In this tutorial, we work with Laya, the open-source decision engine from Convai Innovations that became one of the most-starred machine-learning repositories of September 2026. Laya is a non-autoregressive System 1 model: instead of generating text, a 421-million-parameter encoder reads a piece of text and a set of typed questions, a choice between labels, a score on a scale, or a yes/no, and returns a probability for every option in a single forward pass with zero output tokens. Its pitch is speed and calibrated probabilities, the open answer to TypeSafe’s Jev. Rather than repeat the README’s examples, we put those promises to work on real labelled data with known answers, the banking domain of the CLINC150 intent dataset, and measure what a production router actually gets: zero-shot accuracy against a trained classifier, how much the wording and order of the options matter, how honest the shipped probabilities are, what fitting a temperature on validation data fixes and what it quietly breaks, an abstention gate fitted to an error budget, out-of-scope traffic, a yes/no question that temperature cannot repair, and typed outputs from a pydantic schema.

import os
import sys
import time
import json
import warnings
import traceback
import subprocess
import urllib.request
 
RESULTS = {}
 
 
def banner(title):
    print("\n" + "=" * 78)
    print(title)
    print("=" * 78)
 
 
def section(name):
    def wrap(fn):
        def run(*a, **kw):
            banner(name)
            try:
                out = fn(*a, **kw)
                RESULTS[name] = out if isinstance(out, str) else "ok"
                return out
            except Exception as e:
                RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}"
                print(f"\n[!] {name} did not complete: {type(e).__name__}: {e}")
                traceback.print_exc(limit=3)
                return None
        return run
    return wrap
 
 
banner("1. Install Laya and load the English checkpoint at its reviewed revision")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", "laya==0.3.27"], check=True)
 
import numpy as np
import pandas as pd
import torch
import laya
from laya.calibrate import records_from_labeled
from laya.evals import selective_accuracy, aurc
from laya.common import temp_bucket
 
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
# laya.load() follows the Hub's main branch unless told otherwise. The package ships the commit
# its authors reviewed for each checkpoint; pinning it keeps this notebook's weights fixed.
REVISION = laya.PINNED_REVISIONS["convaiinnovations/laya"]
with warnings.catch_warnings(record=True) as caught:
    warnings.simplefilter("always")
    agent = laya.load("convaiinnovations/laya", device=DEVICE, revision=REVISION)
# On CUDA Laya autocasts to fp16/bf16. Turning that off keeps every device in fp32, so a GPU run
# reproduces the CPU numbers printed below to within floating-point noise.
agent.amp_enabled = False
SHIPPED = (list(agent.temperature), dict(agent.temperature_by_options))
 
n_params = sum(p.numel() for p in agent.model.parameters())
print(f"  laya {laya.__version__}  |  torch {torch.__version__}  |  device {DEVICE}, fp32")
print(f"  checkpoint convaiinnovations/laya @ {REVISION[:7]}  |  {n_params / 1e6:.0f}M parameters"
      f"  |  max_len {agent.cfg['max_len']}, head_max_len {agent.cfg['head_max_len']}")
print("\n  Temperatures shipped with the checkpoint (probabilities = softmax(logits / T)):")
for qt, name in enumerate(["choice", "score", "noul"]):
    print(f"    {name:7s} type-level T = {SHIPPED[0][qt]:.3f}")
for bucket, t in sorted(SHIPPED[1].items()):
    print(f"    {bucket:12s} T = {t:.3f}")
for w in caught:
    if "temperature" in str(w.message):
        print("\n  Warning at load time:\n    " + str(w.message).replace("; ", ";\n    "))
print("\n  T > 1 softens probabilities and T < 1 sharpens them. Remember the choice:11+ row: the")
print("  checkpoint ships 0.10 there, which the loader clamps to 0.5, so any choice question with")
print("  11 or more options gets probabilities SHARPENED by 2x. Step 6 measures what that costs.")

We install the released package, laya 0.3.27, and load the English checkpoint. Two choices here keep the run reproducible. By default laya.load follows the Hugging Face main branch, so we pin the revision the library’s own authors reviewed, which it exposes as laya.PINNED_REVISIONS. And on CUDA Laya autocasts to half precision, so we switch that off to keep every device in fp32 and let a GPU run reproduce the CPU numbers shown here. Printing the checkpoint’s shipped temperatures turns up the first finding before any prediction: the entry for choice questions with eleven or more options is 0.10, outside the valid range, so the loader clamps it to 0.5 and warns. A temperature below one sharpens probabilities, so every answer to a question with that many options will look twice as certain as the raw model is.

TICKET = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
TRIAGE = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
 
 
@section("2. One forward pass, three typed questions, zero output tokens")
def first_decision():
    r = agent.predict(TICKET, TRIAGE)
    a = r["answers"]
    d, u, c = a["department"], a["urgency"], a["churn_risk"]
    print(f"  state: {TICKET!r}\n")
    print(f"  department (choice)  -> {d['choice']!r}   probabilities {d['probabilities']}")
    print(f"  urgency    (score)   -> {u['score']:.2f} on 0..2   level probabilities {u['probabilities']}")
    print(f"  churn_risk (noul)    -> P(yes) = {c['noul']:.3f}")
    print("\n  Two confidence fields, two different quantities:")
    for qid, ans in a.items():
        print(f"    {qid:11s} answer_confidence {ans['answer_confidence']:.3f}   confidence {ans['confidence']:.3f}")
    print("  answer_confidence is the probability of the reported answer, max(p): the number that")
    print("  calibration, the abstention gate and every metric below use. confidence is 1 - normalised")
    print("  entropy, whose scale depends on the number of options. Do not threshold on it.")
    print(f"\n  usage: {r['usage']}")
    print("  output_tokens is always 0: Laya scores the options it is given and never generates text.")
    return f"{d['choice']} / urgency {u['score']:.2f} / P(churn) {c['noul']:.2f} in one pass"
 
 
first_decision()

One call to predict answers three typed questions about a support ticket in a single forward pass: the department as a choice, the urgency as a score from 0 to 2, and the churn risk as a yes/no. The result carries a probability for every option and two confidence fields that are easy to confuse. answer_confidence is the probability of the reported answer, and it is the quantity that calibration, the abstention gate and every metric later in this tutorial use. confidence is one minus the normalized entropy, whose scale depends on how many options a question has. The usage block shows zero output tokens, because Laya scores the options it is given and never generates text.

@section("3. What a pass costs: questions are rows, options are nearly free")
def cost_model():
    def median_ms(q, n=7):
        agent.predict(TICKET, q)
        times = []
        for _ in range(n):
            t0 = time.perf_counter()
            r = agent.predict(TICKET, q)
            times.append(1000 * (time.perf_counter() - t0))
        return float(np.median(times)), r["usage"]["input_tokens"]
 
    print(f"  {'one state, asking ...':34s} {'ms (median of 7)':>16s} {'input_tokens':>13s}")
    rows = {}
    for n in (1, 4, 16):
        q = {f"q{i}": {"type": "noul", "instructions": f"Does the message mention topic number {i}?"} for i in range(n)}
        rows[f"{n} yes/no"] = median_ms(q)
        print(f"  {f'{n:2d} yes/no questions':34s} {rows[f'{n} yes/no'][0]:16.1f} {rows[f'{n} yes/no'][1]:13d}")
    for k in (3, 15, 40):
        q = {"pick": {"type": "choice", "instructions": "Pick one", "criteria": [f"option {i}" for i in range(k)]}}
        rows[f"{k} options"] = median_ms(q)
        print(f"  {f'1 choice question, {k:2d} options':34s} {rows[f'{k} options'][0]:16.1f} {rows[f'{k} options'][1]:13d}")
    print("\n  Every question is encoded as its own (state, question) row, so 16 yes/no questions cost")
    print("  roughly 16 rows. All options of one choice question share a single row and its head")
    print("  budget, so 40 options cost far less than 40 yes/no questions. Design rule: ask one")
    print("  choice question with many options, not many yes/no questions.")
    return (f"16 yes/no {rows['16 yes/no'][0]:.0f} ms vs one 40-option choice "
            f"{rows['40 options'][0]:.0f} ms")
 
 
cost_model()

Before building on Laya, we measure the cost of a forward pass on a single message. Each question becomes its own row, paired with the message, so sixteen yes/no questions take about eight times as long as one. All the options of a choice question share one row and its head budget, so a forty-option choice costs barely twice a three-option one and a quarter of what sixteen yes/no questions do on our CPU. That gives a design rule that shapes everything after it: ask one choice question with many options rather than many yes/no questions.

CARD = json.load(urllib.request.urlopen("https://huggingface.co/api/datasets/clinc/clinc_oos"))["cardData"]
PLUS = next(c for c in CARD["dataset_info"] if c["config_name"] == "plus")
NAMES = {int(k): v for k, v in PLUS["features"][1]["dtype"]["class_label"]["names"].items()}
 
 
def clinc(split):
    df = pd.read_parquet(f"https://huggingface.co/api/datasets/clinc/clinc_oos/parquet/plus/{split}/0.parquet")
    return df.assign(label=df.intent.map(NAMES))
 
 
# The banking domain of CLINC150 (15 intents), with the one-line descriptions a developer would write.
DESCRIBED = {
    "balance": "checking how much money is in an account",
    "transactions": "looking up recent transactions on an account",
    "transfer": "moving money between accounts or to another person",
    "freeze_account": "freezing or locking an account",
    "account_blocked": "an account that is blocked or locked and cannot be used",
    "pay_bill": "paying a bill",
    "bill_balance": "how much is owed on a bill",
    "bill_due": "when a bill is due",
    "interest_rate": "the interest rate on an account",
    "min_payment": "the minimum payment that is due",
    "order_checks": "ordering new checks or a checkbook",
    "pin_change": "changing a PIN",
    "report_fraud": "reporting fraud or suspicious activity",
    "routing": "the bank routing number",
    "spending_history": "how much was spent over a period or on a category",
}
INTENTS = list(DESCRIBED)
ASK = "Which banking request is this?"
 
 
def route(states, criteria, **kw):
    """One choice question over `criteria` for every state; returns choices, confidences, results."""
    res = agent.predict_batch(list(states), {"intent": {"type": "choice", "instructions": ASK, "criteria": criteria}},
                              batch_size=32, **kw)
    return (np.array([r["answers"]["intent"]["choice"] for r in res]),
            np.array([r["answers"]["intent"]["answer_confidence"] for r in res]), res)
 
 
@section("4. Real labelled data: CLINC150 banking, zero-shot vs. a trained classifier")
def zero_shot_vs_trained():
    global train, val, test, BTRAIN, BVAL, BTEST
    train, val, test = clinc("train"), clinc("validation"), clinc("test")
    BTRAIN, BVAL, BTEST = (d[d.label.isin(INTENTS)].reset_index(drop=True) for d in (train, val, test))
    print(f"  CLINC150 'plus': {len(train):,} / {len(val):,} / {len(test):,} train / validation / test queries,"
          f" 150 intents + out-of-scope")
    print(f"  banking domain: {len(INTENTS)} intents; {len(BTRAIN)} train, {len(BVAL)} validation, {len(BTEST)} test queries")
    print(f"  e.g. {BTEST.text[0]!r} -> {BTEST.label[0]}")
 
    t0 = time.time()
    pred, conf, _ = route(BTEST.text, DESCRIBED)
    secs = time.time() - t0
    acc = float((pred == BTEST.label).mean())
    globals()["DESCRIBED_ACC"], globals()["DESCRIBED_PRED"] = acc, pred
    print(f"\n  Laya, zero-shot, 15 options with descriptions: accuracy {acc:.3f}"
          f"   ({secs:.0f}s for {len(BTEST)} queries on {DEVICE})")
 
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression
    print("\n  TF-IDF + logistic regression trained on k labelled queries per intent (5 draws for k < 100):")
    curve = {}
    for k in (1, 3, 10, 30, 100):
        accs = []
        for seed in range(5 if k < 100 else 1):
            sub = BTRAIN.groupby("label", group_keys=False).sample(k, random_state=seed)
            vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
            clf = LogisticRegression(max_iter=3000, C=10).fit(vec.fit_transform(sub.text), sub.label)
            accs.append(float((clf.predict(vec.transform(BTEST.text)) == BTEST.label).mean()))
        curve[k] = float(np.mean(accs))
        print(f"    k = {k:3d}  ({k * len(INTENTS):5,d} labels)   accuracy {curve[k]:.3f}  (sd {np.std(accs):.3f})")
    globals()["CURVE"] = curve
    print("\n  Zero-shot Laya, with nothing but the intent descriptions, lands between what a classic")
    print("  classifier reaches with three and with ten labelled examples per intent. Step 5 shows the")
    print("  descriptions are the weak part.")
    return f"zero-shot {acc:.3f}; TF-IDF needs 10/intent for {curve[10]:.3f}"
 
 
zero_shot_vs_trained()

For real labeled data, we use CLINC150, a public intent-classification benchmark of 150 intents across ten domains plus a set of out-of-scope queries, read straight from the Hugging Face Hub as a parquet file. We take its banking domain, fifteen intents with 100 training, 20 validation, and 30 test queries each, and ask Laya to route the 450 test queries zero-shot, giving it each intent’s name and a one-line description of the kind a developer would write. It reaches 0.804 accuracy with no labeled examples. For scale, a TF-IDF and logistic-regression classifier reaches 0.651 with three labeled queries per intent, 0.848 with ten, and 0.904 with thirty.

@section("5. Criteria wording and option order change the answers")
def wording_and_order():
    t0 = time.time()
    pred_n, conf_n, res_n = route(BTEST.text, INTENTS)
    secs = time.time() - t0
    acc_n = float((pred_n == BTEST.label).mean())
    pred_r, _, _ = route(BTEST.text, INTENTS[::-1])
    acc_r = float((pred_r == BTEST.label).mean())
    flips = float((pred_r != pred_n).mean())
    print(f"  {'criteria':44s} {'accuracy':>8s}")
    print(f"  {'15 names with one-line descriptions (step 4)':44s} {DESCRIBED_ACC:8.3f}")
    tok = {name: agent.predict(BTEST.text[0], {"intent": {"type": "choice", "instructions": ASK, "criteria": c}})
           ["usage"]["input_tokens"] for name, c in (("described", DESCRIBED), ("names", INTENTS))}
    print(f"  {'15 bare intent names':44s} {acc_n:8.3f}   ({secs:.0f}s; {tok['names']} input tokens per query"
          f" vs {tok['described']})")
    print(f"  {'15 bare intent names, order reversed':44s} {acc_r:8.3f}")
    print(f"\n  Reversing the option order changes {flips:.1%} of individual answers, even where the")
    print("  overall accuracy barely moves: the model has a position prior, so fix the order you deploy.")
    changed = pd.DataFrame({"label": BTEST.label, "described": DESCRIBED_PRED, "names": pred_n})
    gained = changed[(changed.names == changed.label) & (changed.de