Doron Tsuberi / Under the hood

Self-improving by design

How to build a self-improving AI product

Four products, four loops, one shape. This is what I would tell a team that wants its AI product to get better with use, drawn from what worked for me in AI-powered products.

Four cases
SuperAgentiChat · DropRenew
OPERATIVY · SPINES

  1. 1

    A unit of output

    One thing you can point at, stamped with the versions that made it.

  2. 2

    A signal from outside

    Outcome, user action, review verdict. The model's self-reflection (a signal from the inside).

  3. 3

    Attribution

    Every signal is traced back through the output to the sources that shaped it.

  4. 4

    A score per source

    Each source that shaped the output is scored by the signals attributed to it: reinforce what helped, suppress what hurt.

  5. 5

    A versioned change

    One layer changes and ships as a new version.

Build them in this order

The five parts

Skipping ahead is how you end up with feedback tables nobody reads.

  1. A unit of output you can point at

    An upload, a verdict, a draft, an answer. Give it an id, and stamp it with the version of everything that produced it: the engine config, the prompt version, the rule versions, the model. Freeze the evidence it was made from. Without the stamp you cannot attribute anything later, and the data you collected is dead weight.

    The stamp is cheap on the day you add it and impossible to add retroactively.

    In practice
    • A decision engine stamps every verdict with its engine version and freezes the factor contributions on the row.
    • A content pipeline snapshots a rule's body before every edit.
    • A research agent records the prompt version on every self-reflection.
  2. A signal from outside the model

    Rank your signals by distance from reality, and know which one you are looking at. Use the cheap signals for volume and the expensive ones as ground truth.

    1. Outcomea sale, a renewal, a conversion. Rare, slow, the only ground truth.
    2. User actionan override with a reason, a thumbs-down, a save. Cheap, plentiful, biased toward whoever bothered.
    3. Human taste at reviewthe operator approving or rejecting a draft. A taste proxy until outcomes arrive.
    4. The model's opinionself-reflection and LLM judges. Unlimited volume, zero authority on its own.

    Never let the model grade itself unchecked.

    In practice
    • A decision engine labels its golden set with outcomes and nothing else.
    • A content pipeline starts from reviewer verdicts, deliberately labelled a taste proxy until funnel results arrive.
    • A research agent verifies each self-reflection claim against the execution trace and marks it confirmed, unverified or contradicted, then calibrates the agent's stated confidence against independent judges. Calibration says when the model's confidence has earned any weight at all.
  3. Attribution back to the thing that produced it

    Joins, not vibes. Record provenance at generation time, never reconstruct it afterwards. Build holdouts in from the start: withhold one rule from a slice of the output, so the rule's effect can be measured instead of assumed.

    A holdout costs nothing at generation time and is the only way to know whether a rule is doing anything.

    In practice
    • A content pipeline writes a usage row for every rule that touched a draft, with the reason it was picked, and withholds one always-on rule from one draft in ten.
    • A decision engine files every override as a hard flip or an adjustment against a specific engine version.
  4. A score per source, computed, not judged

    Attribution says which sources touched a result. The score says, over many results, whether each one helped or hurt: usage and lift per version, outcomes per rule, approval rate per rule, with the holdout as the control. Reinforce above the bar, suppress below it, leave the rest alone.

    Arithmetic over the attributed signals, never a model's impression: the same signals must always give the same score.

    In practice
    • A recognition product scores a pair of items by how often they appear together across all uploads.
    • A content pipeline scores each rule by its approval rate against the holdout.
    • A decision engine scores a candidate configuration by how many past overrides and outcomes it would have got right on a held-out split.
    • A research agent scores each finding by how often it recurs and whether the log confirms it.
  5. A change to one layer, shipped as a new version

    Match the gate to how expensive a wrong change is and how easily it reverses.

    1. Automaticwhere the change is bounded and reversible. No model, no approval, no regret.
    2. Propose onlywhere a wrong change is costly. Beat a bar on held-out data, run in shadow, then a person approves.
    3. Humanwhere taste is the signal. A short list of decisions with the evidence attached.

    Version every change so it can be rolled back. If you cannot say which version made a decision, you cannot learn from the decision.

    In practice
    • A recognition product rebuilds its rankings from counts every night.
    • A decision engine's tuning harness may move two parameters per run, must beat a ship bar on a held-out split, then runs in shadow with a blast-radius report before a person approves.
    • A content pipeline's periodic review puts at most ten decisions in front of the operator, with the evidence attached.
The engine, layer by layerRules and playbooksrules v3Prompts and skillsprompts v7Weights and thresholdsweights v12Corpus and evidencerebuilt nightlyproduceOne unita verdict, an answer, a draftstamped with every layer version:rules v3 · prompts v7 · weights v12shipsThe worldoutcome, override,verdict, votesignal, weeks laterAttributejoin on the stamp,one in ten held outScore, then gateper version, on a held-out splitauto, proposal, or a personprompts v7 becomes v8one layer per signal, the one attribution blamesv7 is kept, so the change can be undoneThe next unit ships stamped rules v3 · prompts v8 · weights v12. Its outcome is compared with the units that carried prompts v7. That is the measurement.
Parts 1 and 3 together, as one loop. Every layer has its own version counter (the numbers are only examples) and each unit is stamped with all of them. Weeks later the outcome arrives and still knows which layer to blame. The change goes back to that layer alone, so any of the four can move, but only one per signal.

What four builds taught me

Five rules

About making the loop turn, not just exist.

  1. 1

    Schedule the last step like a job

    Capture, verification and triage are easy to automate, and they will run whether or not anyone looks. The step that turns findings into a shipped change is the one that needs an owner and a date: a weekly committee, a monthly post-mortem, a review every N drafts. A loop whose last step is "when someone gets to it" is instrumentation.

  2. 2

    No feedback table without a reader in the same change

    Every table that collects a signal ships with the code that consumes it, even if the consumer is only a report on an admin page. A decision engine's override rows feed a production report the day they are created. A content pipeline's rule usage joins to verdicts in the same module that writes it.

  3. 3

    Counting beats learning until you have the data

    A recognition loop can be pure co-occurrence: a million pairs, a handful of thresholds, a nightly rebuild. A name-matching engine that tried embeddings kept its token rules, because coined names do not separate in embedding space. Start with counts and thresholds. Earn the model with data the counts cannot handle.

  4. 4

    Measure the flywheel, not just the product

    Each loop carries one number that says whether it is turning: the match hit rate for a recognition product, the approval rate per rule version for a content pipeline, balanced accuracy and expected-value regret on the held-out split for a decision engine, the calibration gap between stated confidence and judged quality for a research agent. If you cannot state that number, you have a pipeline, not a flywheel.

  5. 5

    Design the reason to contribute

    A loop needs input as much as it needs tuning. A recognition product tells you when a stranger's collection overlaps yours, and opens its API to other agents. A research agent tops up a tester's credits after five pieces of feedback. Ask what the contributor gets back in the same session, and make it visible.

Same shape, four engines

Four loops, side by side

Four kinds of product, the same five parts. For each: the forward pass, the signal that comes back, what it scores, what it updates, and the gate it passes.

A research agentunit of output: one answerQuestionData toolsPrompts, skillsAnswergatetriage score, then reviewsignalvotes, independent judges, self-reports checked against the logscoresjudge scores per prompt and skill version; findings by recurrenceupdatesprompt, tool and skill versionssame five parts; the signal and the gate differA decision engineunit of output: a keep-or-drop verdictItemScoring rulesReferee modelVerdictgateship bar, shadow run, approvalsignaluser overrides with reasons; real outcomes 90 days laterscoresa candidate config by the overrides and outcomes it gets rightupdatesrule weights, two per run, new engine versionsame five parts; the signal and the gate differA content pipelineunit of output: a draftTopicPlaybookDraftReview queuegatereview pass, ten decisions maxsignalreviewer verdicts; funnel events; one in ten held outscoreseach rule by its approval rate against the holdoutupdatesplaybook entries kept, revised or retiredsame five parts; the signal and the gate differA recognition productunit of output: an approved uploadPhotoVision modelCorpus matchPair countsgateautomatic, thresholds onlysignalapproved uploads, matches between users, viewsscoreseach pair of items by how often they appear togetherupdatescounts and rankings, rebuilt nightlysame five parts; the signal and the gate differ
The signal moves from cheap to expensive, and the gate from automatic to a person, as a wrong change gets dearer.