In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the paper’s idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and its edit history, and then plug a simulated agent into RRSI’s own Domain interface. Because we built the simulated environment ourselves, we know the true effect of every edit, which lets us audit RRSI’s decisions against ground truth and compare them with an unregularized search that simply keeps whatever scores highest.
import os
import sys
import json
import math
import copy
import random
import tempfile
import textwrap
import traceback
import subprocess
import statistics as st
from pathlib import Path
RESULTS = {}
def banner(title):
print("\n" + "=" * 78)
print(title)
print("=" * 78)
def section(name):
def wrap(fn):
def run(*a, **kw):
banner(name)
try:
out = fn(*a, **kw)
RESULTS[name] = out if isinstance(out, str) else "ok"
return out
except Exception as e:
RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}"
print(f"\n[!] {name} did not complete: {type(e).__name__}: {e}")
traceback.print_exc(limit=3)
return None
return run
return wrap
banner("1. Install RRSI and map the paper onto the code")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", "git+https://github.com/google-research/rrsi.git@be50316e1db05914068a973f322770ef08ed7ba1"], check=True)
from importlib.metadata import version
from rrsi.config import RRSIConfig
from rrsi.evaluate import TaskResult, EvalResult, aggregate, evaluate
from rrsi.calibrate import calibrate
from rrsi.selection import Candidate, cost_rule, judge, select_round
from rrsi.schedule import edit_budget, budget_table
from rrsi.history import History, stall_flag, exploration
from rrsi.components import K, K_STR, normalize, novelty
from rrsi.critic import precheck, review
from rrsi.domain import Domain
print(f" rrsi {version('rrsi')} | anthropic {version('anthropic')} | Python {sys.version.split()[0]}")
print("\n RRSI evolves an agent's HARNESS (prompts, tools, memory, control flow, sub-agents) around a")
print(" frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker")
print(" benchmarks. The part that decides which edits to KEEP is plain Python, and that is what we drive:")
for symbol, where in [
("S_hat, C_hat Eq. (estimate)", "rrsi.evaluate.aggregate"),
("delta noise band", "rrsi.calibrate.calibrate"),
("Algorithm 2 floor + cost rule", "rrsi.selection.judge / select_round"),
("b_t Eq. (anneal)", "rrsi.schedule.edit_budget"),
("Critic leakage screen", "rrsi.critic.precheck / review"),
("L_t, g_t, B_t history, yield, prune", "rrsi.history.History"),
("sigma_t, U_t stall + exploration", "rrsi.history.stall_flag / exploration"),
("nu structural novelty", "rrsi.components.novelty"),
]:
print(f" {symbol:40s} -> {where}")
CFG = RRSIConfig()
print(f"\n paper defaults: T={CFG.T} rounds, k={CFG.k} trials/task, m={CFG.m} candidates/round,"
f" b in [{CFG.b_min},{CFG.b_max}]")
print(f" beta0={CFG.beta0} beta1={CFG.beta1} w_s={CFG.w_s} w_c={CFG.w_c} w_n={CFG.w_n}"
f" delta_z={CFG.delta_z}")
print("\n Nothing below needs an API key, a GPU or a dataset download.")
We install RRSI from the google-research repository, pinned to the commit this notebook was written against, since the package is not on PyPI. Its only dependency is the Anthropic client, which the search roles use to call Claude and which we never exercise. We then print the mapping the repository itself documents between the paper’s symbols and the functions that implement them: the empirical score and cost estimate in evaluate, the noise band in calibrate, Algorithm 2 in selection, the annealed edit budget in schedule, the leakage screen in critic, and the edit history with its yield, prune, stall and exploration summaries in history. RRSIConfig holds the paper’s hyperparameters, and every function below receives it exactly as the real loop does.
@section("2. Evaluate(H): a score and a cost, and why a crash counts as zero")
def estimator():
base = {
"task_000": TaskResult(rewards=[1, 1], tokens=[11_800, 12_400]),
"task_001": TaskResult(rewards=[1, 0], tokens=[15_100, 14_600]),
"task_002": TaskResult(rewards=[0, 0], tokens=[21_000, 19_500]),
}
ev = aggregate("H0", 2, base)
print(f" three tasks x k=2 trials -> S_hat = {ev.S:.3f} C_hat = {ev.C:,.0f} tokens/trial"
f" ({ev.n_expected} trials expected, {ev.missing} missing)")
crashy = dict(base)
crashy["task_002"] = TaskResult(rewards=[0.0, 0.0], tokens=[None, None], missing=2)
ev_crash = aggregate("crashy", 2, crashy)
dropped = {t: r for t, r in crashy.items() if not r.missing}
naive = sum(sum(r.rewards) for r in dropped.values()) / sum(len(r.rewards) for r in dropped.values())
print("\n A candidate crashes on the hardest task instead of failing it:")
print(f" an estimator that drops missing trials reports {naive:.3f} <- looks like a gain")
print(f" RRSI's aggregate (missing = 0, full denominator) {ev_crash.S:.3f} <- no reward for crashing")
print(f" ...and C_hat uses only recorded token counts: {ev_crash.C:,.0f}")
rubric = {"memo": TaskResult(rewards=[0.5, 1.0], weights=[10, 10]),
"brief": TaskResult(rewards=[0.0, 0.0], weights=[90, 90])}
ev_w = aggregate("rubric", 2, rubric)
print("\n Weighted rewards (Harvey LAB style: weight = number of rubric criteria):")
print(f" mean of per-task means = {st.mean(r.mean for r in rubric.values()):.3f}"
f" vs RRSI's S_hat = {ev_w.S:.3f} (fraction of all criteria passed)")
return f"crash scored {ev_crash.S:.3f} under RRSI vs {naive:.3f} if dropped"
estimator()
RRSI measures two numbers per harness: S, the reward averaged over every trial of every task, and C, the mean policy tokens per trial. TaskResult records one task’s trials and aggregates them. The detail worth copying into any agent evaluation is how missing trials are handled. When a candidate crashes on the hardest task, an estimator that drops the missing trials reports 0.750 and makes the crash look like an improvement. At the same time, RRSI counts each missing trial as zero reward with the full denominator and reports the same 0.500 as before, so a candidate cannot look better by destroying the trials it finds hard. Weighted rewards cover rubric-graded suites such as Harvey LAB, where S becomes the fraction of all criteria passed rather than the mean of per-task means.
FAMILIES = ["parse", "search", "edit", "test"]
H0 = {"skill": {f: 0.0 for f in FAMILIES}, "memo": [], "cost": 1.0}
def make_world(seed, n_evolve, n_heldout=80):
"""Tasks with a family and a difficulty. Evolve ids are task_NNN, held-out ids held_NNN."""
r = random.Random(seed)
world = {f"task_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_evolve)}
world.update({f"held_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_heldout)})
return world
def p_success(h, world, task):
"""The frozen policy: a logistic in harness skill minus task difficulty, or 0.97 if memorised."""
if task in h["memo"]:
return 0.97
family, difficulty = world[task]
return 1 / (1 + math.exp(-(0.3 + h["skill"][family] - difficulty)))
def run_trials(h, world, ids, k, rng):
return {t: TaskResult(rewards=[float(rng.random() < p_success(h, world, t)) for _ in range(k)],
tokens=[round(12_000 * h["cost"] * math.exp(rng.gauss(0, 0.08))) for _ in range(k)])
for t in ids}
def true_score(h, world, ids):
return sum(p_success(h, world, t) for t in ids) / len(ids)
def calibrated_delta(world, k, seed, repeats=3):
"""delta from `repeats` independent evaluations of the SAME harness H0."""
ids = [t for t in world if t.startswith("task_")]
rng = random.Random(10_000 + seed)
evals = [aggregate(f"base{j}", k, run_trials(H0, world, ids, k, rng)) for j in range(repeats)]
return calibrate(evals, z=CFG.delta_z, reps=200), evals
@section("3. The noise band: how far one harness's score moves between two evaluations")
def noise_band():
world = make_world(0, 40)
ids = [t for t in world if t.startswith("task_")]
rng = random.Random(1)
scores = [aggregate(f"H0#{j}", 2, run_trials(H0, world, ids, 2, rng)).S for j in range(6)]
print(f" H0 evaluated 6 times on 40 tasks x k=2: {[round(s, 3) for s in scores]}")
print(f" true expected score {true_score(H0, world, ids):.3f}; spread of the estimates"
f" {max(scores) - min(scores):.3f}. Nothing about the harness changed.")
print(f"\n calibrate() turns repeated evaluations into delta = z * sd(null dS), z = {CFG.delta_z}:")
print(f" {'evaluator':>12s} {'trials':>7s} {'delta':>8s} method")
deltas = {}
for n, k in [(40, 2), (100, 4), (400, 8)]:
cal, _ = calibrated_delta(make_world(0, n), k, seed=0)
deltas[f"{n}x{k}"] = cal["delta"]
print(f" {f'{n} x k={k}':>12s} {n * k:>7d} {cal['delta']:>8.4f} {cal['method']}")
print("\n The paper's calibrated bands are 0.017 (coding), 0.004 (workspace) and 0.020 (engineering).")
print(" A gain smaller than delta is indistinguishable from re-running the same harness, and")
print(" Algorithm 2 treats it that way. Remember the first row: it becomes the lesson of step 11.")
return "delta " + ", ".join(f"{k}={v:.3f}" for k, v in deltas.items())
noise_band()
Before any rule can separate a real gain from luck, it needs to know how far one harness’s score moves on its own. We build a small simulated agent, whose success on each task is a logistic function of harness skill minus task difficulty, and evaluate the unchanged starting harness six times on forty tasks with two trials each: the scores spread by 0.113 although nothing changed. calibrate turns repeated evaluations of the same harness into delta, twice the standard deviation of the difference between two runs. With 80 trials delta is about 0.108; with 3,200 trials it falls to about 0.013, in the range the paper reports for its instances (0.004 to 0.020). For selection purposes, any gain smaller than delta is indistinguishable from re-running the same harness.
def ev_at(S, C, job, n=200):
"""An EvalResult with exactly score S (in steps of 1/n) and C tokens per trial."""
hits = round(S * n)
return aggregate(job, 1, {f"task_{i:03d}": TaskResult(rewards=[1.0 if i < hits else 0.0], tokens=[C])
for i in range(n)})
INCUMBENT = ev_at(0.630, 10_000, "incumbent")
S_STAR, DELTA = 0.640, 0.020
@section("4. Algorithm 2, branch by branch: the floor, the cost rule and the noise band")
def algorithm_2():
print(f" incumbent S = {INCUMBENT.S:.3f}, C = {INCUMBENT.C:,.0f} best ever S* = {S_STAR} delta = {DELTA}")
print(f" floor = S* - delta = {S_STAR - DELTA:.3f}\n")
cases = [
("A slipped below the floor", 0.600, 10_000, ["prompt"], None),
("B real gain, pays for its tokens", 0.690, 12_000, ["prompt"], None),
("C real gain, far too expensive", 0.660, 25_000, ["subagent"], None),
("D slightly WORSE but cheaper", 0.625, 8_000, ["context_mgmt"], None),
("E in the band but costlier", 0.640, 11_500, ["prompt"], None),
("F in the band, new sub-agent", 0.630, 10_000, ["subagent"], None),
("F' in the band, prompt tweak", 0.630, 10_000, ["prompt"], None),
("G like B, breaks a domain guard", 0.690, 12_000, ["prompt"], ["valid-output rate fell"]),
]
rows = []
for label, S, C, comps, guards in cases:
cand = Candidate(label[:2].strip(), [{"id": "C1", "component": c} for c in comps], ev=ev_at(S, C, label))
dec = judge(cand, INCUMBENT, S_STAR, DELTA, CFG, incumbent_counts={}, guards=guards)
rows.append((label, dec))
print(f" {label:36s} S={S:.3f} C={C:>6,} {'ADMIT ' if dec.admissible else 'reject'}")
print(textwrap.indent(textwrap.fill(dec.reason, 88), " " * 6))
print("\n Three rules, in the order RRSI applies them:")
print(" 1. never fall below the best score ever seen, minus the noise band (A)")
print(" 2. a gain bigger than delta must pay for any extra tokens: dC <= 0.10 + 40*dS (B, C)")
print(" 3. inside the band scores are a tie, so prefer the cheaper harness, and let a")
print(" never-tried structural component break the tie (D, E, F, F')")
print(" D is the surprising one: RRSI admits a harness that scored LOWER than the incumbent.")
return f"{sum(d.admissible for _, d in rows)}/{len(rows)} admissible; D admitted at dS={rows[3][1].delta_S:+.3f}"
algorithm_2()
Algorithm 2 is implemented as pure functions, so we can hand it candidates and read its reasons verbatim. We fix an incumbent at S 0.630 and 10,000 tokens, a best-ever score of 0.640 and delta 0.020, and send eight candidates through judge. A candidate below the floor, the best score ever seen minus delta, is rejected outright. A gain larger than delta must pay for its extra tokens under the rule that the relative cost change stays below 0.10 plus 40 times the gain; a +6 point gain at +20% tokens is admitted, and a +3 point gain at +150% tokens is not. Inside the band, scores are treated as a tie and a shaped score of 100 times the gain minus 15 times the cost change, plus a small bonus for a never-accepted structural component, decides; that is how RRSI admits candidate D, which scored lower than the incumbent but costs 20% fewer tokens, and how a new sub-agent breaks a tie that an equally scoring prompt tweak does not. A domain guard vetoes regardless of score.
@section("5. One round of selection: the highest score does not always win")
def one_round():
c_hi = Candidate("C", [{"id": "C1", "component": "subagent"}], ev=ev_at(0.660, 25_000, "C"))
d_lo = Candidate("D", [{"id": "C1", "component": "context_mgmt"}], ev=ev_at(0.625, 8_000, "D"))
leak = Candidate("X", [{"id": "C1", "component": "memory"}], gate_failure="critic_reject")
winner, decisions = select_round([c_hi, d_lo, leak], INCUMBENT, S_STAR, DELTA, CFG, incumbent_counts={})
for d in decisions:
print(f" {d.variant}: {'admissible' if d.admissible else 'rejected '} {d.reason[:92]}")
new_star = max(S_STAR, winner.ev.S)
print(f"\n winner: {winner.variant} (S {winner.ev.S:.3f}, C {winner.ev.C:,.0f})")
print(f" the new incumbent scores {winner.ev.S - INCUMBENT.S:+.3f} vs the old one and costs"
f" {(winner.ev.C - INCUMBENT.C) / INCUMBENT.C:+.0%} tokens; S* stays {new_star:.3f}")
print("\n C scored highest and still lost: its +3pp does not pay for +150% tokens. D moved the")
print(" incumbent DOWN inside the noise band because it is 20% cheaper. Because S* only ever")
print(" rises, the floor never follows the incumbent down, so a chain of 'cheaper but slightly")
print(" worse' swaps cannot walk the score away. X never reached evaluation at all.")
return f"winner {win