Evolution · Self-improving VIs Crucible & Axon

The models that improve themselves

Close the loop.

A VI that ships is never finished. Crucible turns your production traffic into evaluations, mutates candidate models against them, and promotes only the ones that measurably win — the same selection pressure from our lab, pointed at your metrics.

01 — CrucibleTrace → Evaluate → Mutate → Promote

From a live trace to a better model, with no one in the loop.

The manual version of this is a team of ML engineers reading logs and hand-writing evals. Crucible runs it continuously, and shows its work at every step.

TRACE EVALUATE MUTATE PROMOTE autonomous NO HUMAN IN THE INNER LOOP
1 · Trace
Every production call — prompt, tools, retrieval, output, and user reaction — is captured as a structured trace. Recurring behaviours are clustered into a map you can inspect.
2 · Evaluate
From those clusters and real feedback, Crucible writes evaluations automatically and generates production-realistic synthetic data to cover the edge cases your traffic hasn't hit yet.
3 · Mutate
Candidate variants — new prompts, tool schemas, routing, or fine-tunes — are spawned and scored against the eval set. Regressions are caught before anything reaches a user.
4 · Promote
Only variants that beat the incumbent on your metrics are promoted, carrying a full genealogy of what changed and why. The loop restarts on the new baseline.
02 — How it's builtEvaluation, then automation

An eval engine, wearing an autopilot.

Two ideas stacked: rigorous evaluation of agent behaviour, and an operations layer that acts on it without waiting for a human. One makes the model knowable; the other makes improvement continuous.

A

Ingest traces, group recurring patterns into behaviours a person can actually read.

Raw logs are useless at scale. Crucible normalizes every trace and clusters them into a behaviour map — "the agent over-apologizes on refunds," "retrieval misses when the query is a date." Each behaviour is inspectable, countable, and trend-able.

This is the same move our world models make on sensory data: compress a flood of observations into a small set of structures you can reason about.

B

Turn behaviours and feedback into evaluations that catch regressions before deploy.

For every behaviour worth keeping — or killing — Crucible drafts an eval: an input, an expected property, and a grader. Thumbs-down, escalations, and corrections become test cases automatically.

The eval suite grows with your product instead of rotting behind it. Nothing gets promoted that quietly breaks something a user already relied on.

C

Generate production-realistic synthetic data for the cases you haven't seen yet.

Real traffic under-samples the dangerous tail. Crucible synthesizes plausible edge cases — rare intents, adversarial phrasings, malformed inputs — so a variant is stress-tested against the future, not just the past.

Coverage becomes a number you can watch climb, not a hope.

D

Build an operational world model that links deploys to metrics, alerts, and owners.

To automate safely, the system needs context: which deploy touched which service, what "normal" looks like, who to route a change to. Crucible maintains a live model of your stack — connected to your CI, dashboards, and incident tools — updated in real time.

It's the difference between an autopilot with a map of the terrain and one flying blind.

E

Deploy patrol and triage agents that watch for drift and turn noise into structured issues.

A patrol agent proactively hunts for model drift, latency creep, and quality regressions. A triage agent collapses noisy signals into a single, structured issue with a proposed fix — and, where policy allows, opens the pull request.

Engineers set direction and approve. The agents do the tracing, the reproduction, and the first draft of the fix.

F

Route every promotion through your existing approvals. Automation proposes; people dispose.

Changes flow through the pipelines you already trust — pull requests, staged rollouts, and sign-off in chat. The inner loop is autonomous; the act of shipping to users is gated by a human, on purpose.

Fast where speed is safe, deliberate where it isn't.

03 — AxonUnified multi-agent runtime

One interface over every model you run.

Crucible improves a model. Axon is where models live and work together — a single operations layer that consolidates data, orchestrates agents, and automates the standard-operating-procedure work of a whole organization.

Any LLM, one API

Route any prompt to any model — ours, open-weight, or a frontier vendor — and switch providers without touching your code. Auto-detect complexity and spend tokens where they matter.

Grounded retrieval

Unified enterprise search with connectors for your document stores and systems of record. Citations down to the exact source, with permissions mirrored automatically.

SOP-driven workflows

Compose sequential and parallel agent chains on a visual canvas or in code. Turn a standard operating procedure into an auditable, repeatable, VI-run process.

Full observability

Visualize the reasoning chain from intent to execution. Real-time audit trails, policy-violation alerts, and explainability built for regulated work.

Voice & channel-native

Deploy the same VI to chat, voice, and IVR with real-time transcription and clean hand-off to a human — full context transferred, nothing lost.

Runs anywhere

On our sovereign compute or inside your own cloud and private landscape. Prototype to production in weeks, replacing a stack of point solutions.

Axon orchestrates VIs. It does not grant them autonomy: every workflow is scoped, logged, and reversible, and every consequential action can require human approval.

04 — Where this goesTool today, frontier tomorrow

Automation now. Evolution, carefully.

  1. Phase 1

    Assisted

    Crucible surfaces behaviours and drafts evals; engineers approve every change. Available today.

  2. Phase 2 · now

    Automated inner loop

    Patrol and triage agents propose fixes and open PRs. Promotion stays human-gated.

  3. Phase 3

    Population training

    Whole populations of variants evolve against your evals in our sandbox, not just single candidates.

  4. Phase 4

    Self-directed research

    Models that propose their own fitness functions — strictly a lab frontier, never a shipped VI.