rathvan

The case

Generation got cheap. Knowing it is right did not.

A coding agent will write you a service this afternoon. It will not tell you whether the tenancy predicate is enforced, whether the thing you asked for already exists, or which of your decisions cannot be undone. Rathvan is built around that second half — the part that stayed expensive, and the part that decides whether any of this is safe to run unattended.

For the engineer evaluating it

Four things a diff review cannot give you

Every one of these is a property of the pipeline, not a feature you configure.

Deterministic gates

Programs judge first, you judge last

Checkers read the requirements, the design and the tests before code exists. They are string comparisons against a source tree, not a second model grading the first — every finding can be shown to be true, and the same input always produces the same verdict.

Repair before review

The system passes its own gate first

Findings go back to the generator in a phase that may still act on them, bounded at two attempts, keeping the better draft rather than the most recent. What reaches you is what a program could not decide.

Provenance

A human edit is visible as one

Artifacts version rather than overwrite, and a person's edit records the person, not a model id. Two months later a reviewer can see which lines a human wrote — the difference between reviewing a document and auditing one.

Reuse as arithmetic

"Does this exist already" has an answer

A catalogue of 139 capabilities across 14 domains, and a checker that reads a design's claims against the kernel's declared types. A type that does not exist is refused by name.

1825kernel tests
none skipped
109migrations
additive only
139capabilities
14 domains
13pipeline stages
3 human gates
7tracks
rigour by stakes

For the enterprise

Adoption without a migration

The usual objection to a platform is that it is somewhere you have to move to. This one is a kernel you adopt, and the difference is visible in the signature.

The question a platform team asksThe answer here
What do we have to supply? Four ports — persistence, a model chain, source control, a feature gate. No framework, no database and no HTTP inside the kernel.
How do we know multi-tenancy is right? The tenant is a parameter on every store method, not an ambient context — so an adapter that is single-tenant does not compile. Correct by construction rather than by having read our documentation.
Where does our specification go? Wherever you point it. The model chain includes adapters for models served from your own hardware, so a build that answered data residency: yes can generate without a byte leaving your infrastructure. A user's own key is supported for the same reason.
What does it do to our database? Additive migrations only. Row-level security enabled and forced on every table holding user data, asserted in the first migration and held by a test.
Can it reach production on its own? No — and not as a setting. The state machine has no production transition, because the ways a deploy fails are invisible to every signal a pipeline can read.
What does an audit look like? Every transition names an approver and is recorded; every generation is metered by vendor, model and token; a forced gate keeps the reason forever.

Why this compounds

The part a model cannot generate for you

The code here is the smaller asset. What accumulates is harder to copy and gets more valuable with every build that runs through it.

The method

Written down, machine-readable, and under test

13 stages and 3 gates in a manifest the console renders and the kernel validates, beside a dated log of the ways a check reported success while the work was broken. That log is operational experience, not documentation.

The catalogue

139 capabilities, ordered by retrofit cost

Not a feature list — a map with lifecycle and readiness per capability, arranged by what it costs to add later rather than by architectural tidiness. Reuse decisions are looked up against it and then verified.

The gates

Each one exists because something got through

Every checker in this platform was written after a specific failure, and carries the evidence. A competitor can copy the idea of a quality gate; the value is in knowing which ones to have, and that is bought one incident at a time.

The proof

It is built by itself

Every enhancement to this platform goes through this platform's own pipeline. That is the only honest way to find out whether a process is tolerable — and it is how the gates get found wrong, which is the point.

See it end to end

From an empty workspace to a live application

Thirteen steps, in the order an enterprise actually meets them — set up the workspace, add people, connect your own accounts, describe what you need, approve at every gate, watch it deploy to your cloud or your own machine, open the running application, and read what it cost. Each screen mirrors a surface that exists. Nothing here calls a builder; it is a walkthrough, and it says so at the top of every step.

Walkthrough · not a live builder

Three screens are worth pausing on, and a code generator has no equivalent for any of them. At Design review the checkers have already run and already sent one draft back, so what reaches a person is what a program could not decide. At Live the data tab shows the tenancy predicate that fences the rows — forced, not merely enabled, because otherwise the table owner bypasses every policy. And at Receipt every call is attributed to a vendor and a model, because a build you cannot account for is one you cannot put in front of a finance team.

A captured run · 26 August 2026

What it looks like when the gate says no

This is a real build on the production builder, driven from the command line and copied here unedited. It is shown rather than a clean pass because a gate that only ever agrees with the generator is not a gate, and this one refused.

$ rathvan build "Let an advisor see which filings are overdue for a client"
  #131 started — PRD_REVIEW

$ rathvan build approve 131
  #131 approved — now DESIGN_REVIEW

$ rathvan build approve 131          # tests, build spec, a branch
  #131 approved — now BUILDING

$ rathvan build approve 131          # through the quality gate
  #131 approved — now GATES

$ rathvan build approve 131          # open the pull request
  Refused: 2 blocking finding(s):
    ReuseLint · com.rathvan.mcp.actions.OverdueFilingAction
    ReuseLint · com.rathvan.mcp.McpToolCatalogue.toolsFor
  — this phase cannot regenerate anything (it reads verdicts; it does not call
    models), so the two ways on are: fork the build and fix the artifact in a
    phase that may, or force past this with a written reason that is kept.

$ rathvan build status 131
  #131  GATES
    BUILD_SPEC v1
    DESIGN v1
    DESIGN v2      # repaired before a reviewer saw it
    DESIGN v3      # repaired again, then the loop stopped
    PRD v1
    TESTS v1
What was caught

A type and a method that do not exist

The design claimed to use OverdueFilingAction and to call McpToolCatalogue.toolsFor. Neither is in the kernel. This is a string comparison against a source tree, not a second model's opinion — every finding can be shown to be true, and it names the thing by name.

What was tried first

Two repairs, before a person was asked

DESIGN v1 → v2 → v3. The checkers ran at generation time and handed their findings back twice. A model shown a deterministic finding twice is telling you it cannot act on it — so the loop stops rather than burning a third pass.

What happens now

Two honest ways forward, and no third

Fork the build and fix the design where a model may still be called, or override the gate with a written reason that is kept forever, with your name on it. There is no button that makes the finding go away quietly.

An earlier build the same day — #130 — passed the same gate with zero blocking findings and 8 of 8 acceptance criteria covered by tests, and opened a pull request. The difference between the two is what the generator happened to write, which is the point: the verdict is a property of the artifact, not of the day.

The vision

Nobody should start from an empty directory again

Every company building software rebuilds the same floor: identity, tenancy, audit, consent, money in minor units, feature flags. It is rebuilt because it is invisible until it is wrong, and by then it cannot be changed. The product a company actually sells sits on top of that floor and is the cheapest part of the stack to change.

The end state is that the floor stops being written at all. A team describes what they need, the platform composes what exists, generates only what genuinely is new, and a human approves the decisions that cannot be undone. What ships is a repository that builds, with the compliance regime already in it — not a starting point somebody has to finish.

That is not a claim about better code generation. It is a claim about where the expensive decisions get made, and about moving them from month nine to week one, where they are still cheap.

The path

Three horizons, each with something that can fail

A roadmap where nothing can be falsified is a wish list. Each horizon below states the thing that has to become true, and the next one does not start until it has.

  1. Now — prove it on somebody else's work

    A team that did not build it, builds with it

    Everything described on this page runs, and none of it has met the one test that matters. The work here is adoption friction, not features: the last mile between an engine that generates correct artifacts and a person who can operate it without us in the room.

    Passes when a team outside this company takes a product from a sentence to a merged pull request, unaided, and the gates they meet are ones they agree with rather than ones they force past.

  2. Next — the platform edits what already exists

    Brownfield, not just greenfield

    Scaffolding a new repository is the easier half and the smaller market. Most engineering is changing a codebase that already exists, under tests that already pass, without breaking what is around it. The reading and the commit path work; the generation of a bounded change is the frontier.

    Passes when a change to an existing repository — one nobody here wrote — goes through the full pipeline and is merged on its own evidence.

  3. Then — the method becomes the product

    Verification other people can adopt

    The durable asset is not the generator, it is the set of checks and the order they run in. A platform team should be able to bring their own standards, register their own gates, and hold their own agents to them — with the catalogue and the failure log as the starting position rather than something they assemble over years.

    Passes when an organisation runs gates they wrote themselves, on builds we never see, and the audit trail satisfies their reviewers rather than ours.

Upcoming

What is being built next, and why in that order

Each of these is tracked in the backlog with a condition that says when it is finished. They are sequenced by what unblocks what, not by what demos best.

Next

Operating it without us in the room

The engine generates correct artifacts today. The work ahead is the last mile — the difference between a pipeline that produces the right documents and a team that can run it unaided, meeting gates they agree with rather than force past.

Next

Changing a repository that already exists

Reading a repository, proposing a change and committing it all work now. Generating a bounded change to a large existing file is the frontier, and it is the larger market — most engineering is not greenfield.

Upcoming

Bring your own capacity, and your own gates

Generation already runs on your key, in your region, with adapters for models served from your own hardware. Ahead of that: registering your own checkers, so a platform team holds its agents to standards it wrote rather than ones we did.

Upcoming

Prompt caching across every vendor

The chain measures every call and reports cached share per build. Extending cache-aware requests across all adapters is a direct reduction in what a build costs, and the meter is already there to prove it either way.

The platform's own standard is that a check which cannot say what it could not see is worth less than one that can. The roadmap is held to the same rule: each item above has a written condition for being struck, so “upcoming” has an end.