Work/How it learnsTerms explained in the pattern
SuperAgentiChat: the agent files a report on itself after every answer.
SuperAgentiChat answers financial research questions by running tools over filings and market data. After every answer the agent writes a private note on what went wrong. The note is checked against what actually happened, independent judges score the answer, and the findings become changes to its prompts, tools and skills.
Under the hood
The loop, stage by stage
Each stage is a mechanism you could build yourself: what is captured, how it is checked, and what it changes.
-
The loop runs
A question fans out over 40+ tools, from filings and market data to news, macro series and prediction markets, plus 70+ built-in skills, from valuation models to screeners. Skills are data, not code: they are loaded per user and scope from a table, and nothing is hardcoded in the prompt.
How it is built- Tools are registered in one place. A handful more are attached per request, such as reading the user's own files and the self-critique tool.
- Skills live in a capabilities table with entry types for tools, skills, sources and prompts, merged from user, public and system scopes.
- Several model providers, each paired with a cheaper router model for judging and digests.
-
The hidden reflection
Exactly once per turn, alongside the final answer, the agent calls a diagnostic tool and fills in a structured self-critique. The user never sees it, it costs the answer nothing, and it is stripped from history before the next turn.
How it is built- The self-critique tool is versioned and described to the model as an internal diagnostic, never to be discussed with users.
- Fields: answer quality, confidence, hallucination risk and areas, intent clarity, assumptions, tool issues by type, web fallbacks with the source that should have existed, data gaps, and typed suggestions: new tool, tool improvement, skill improvement, prompt, step budget.
- The step limit is raised by the reflection's own cost, so it never steals a step from the answer. If the model skips the call, a missed row is filed anyway.
- Stored in its own table with the columns denormalised for querying, plus provider, model and star flags.
-
The reflection is not trusted
A self-assessment is only useful if it is checked. Each claim is verified against the real trace and marked confirmed, unverified or contradicted. Separately, up to seven judges score the answer itself, and users can vote it up or down.
How it is built- Verification produces lines like “Contradicted: the agent claims the step budget was insufficient, but used only a fraction of its steps.”
- Evaluation runs deterministic checks plus judges for numerical accuracy, analytical quality, compliance tone, temporal consistency and tool-use correctness, with cross-source reconciliation and a calculation chain when relevant, into one composite score.
- Thumbs up or down with a comment are stored per message. Testers earn credit top-ups after five submissions.
- Calibration compares the agent's stated confidence with judge scores. That gap is the signal for when the agent's own confidence can be trusted.
-
Everything is traced in-house
There is no third-party tracing vendor. Every request, tool call and error lands in the product's own tables, which is what makes the reflection verifiable and the judges auditable.
How it is built- Request traces, tool invocations and error events are the ground truth that verification checks claims against.
- Reflections carry admin and eval star flags, so notable cases can be pulled into review sets.
- A security directive in the system prompt treats any user question about the reflection as prompt probing.
-
Triage
Hundreds of reflections are noise-filtered, near-duplicates merged, and each finding scored on frequency, impact and novelty against the effort to fix it. An admin dashboard turns a batch into a digest; a documented cycle turns the digest into action items.
How it is built- The triage pipeline: noise filter, near-duplicate merge, then a weighted score of frequency, impact and novelty divided by effort.
- Digests are saved and read on an admin dashboard.
- A documented review cycle turns each digest into numbered action items. Fixes are applied in order: prompt, then tool or routing, then architecture.
-
New tools and skills come through a scouting process
Suggestions typed as new tool or new skill do not go straight into the product. They become candidates for a scouting process in which external models survey what a financial analyst would need, and the winners are implemented as skills or tools with a written guide.
How it is built- Reflection suggestions typed as new tool or new skill become scout candidates; recurring failures become tracked issues.
- The scouting process has catalogued 135 ideas over fifteen sprints, 59 of them implemented.
- A new skill is a migration against the capabilities table, following a written guide. Users, the public scope and the system scope can each carry their own.
-
What has changed because of it
The loop has already rewritten the system prompt more than once. The clearest case came from a failure-review digest rather than from reflections.
How it is built- A failure classifier found an 82% failure rate on one class of question, most of it hallucinated numbers and scale mismatches. Result: a new system prompt version with a rule to ground every number in tool output.
- One incident: the model called its own reflection tool eleven times in one turn. Fix: the tool now answers “recorded, do not call again”, shipped in the next prompt version.