Let Agents Write Code. Make Them Prove It's Right.

Jan 13, 2026
Steven Klaiber-Noble
Software Engineer
Deterministic Boundary Programming: constraining AI with rigid boundaries

AI lets you write more code than ever, but you aren't shipping much more of it.

Your coding agent produces a lot of plausible-looking code, but much of it never reaches production. One reason is silent failure, where code comes back marked done without being functional. In one documented case, an agent passed in nonexistent data, saw an error, and still considered the task complete. The other is slop, the gradual drift from small inconsistencies that no one catches, until you can't trust the codebase anymore.

Both problems come from the same place. The agent has nothing deterministic to check its work against. The fix is structural. Let the agent write every layer, but arrange the layers so each one is forced to obey the one beneath it, down to rules it can't change. We call this Deterministic Boundary Programming. It works on any stack, and FRAGMENT is built around it, so the agent has to work within these boundaries and produce code that does what you intended.


Why Code Piles Up But Doesn't Ship

The mistake is letting your agent run with nothing deterministic to push against.

You prompt it to "build me a payments flow," and it does whatever makes the code run. It invents types and makes up field names that sound right. The code compiles, it looks plausible, and the agent calls it done. Then you run it against real data and the payment never goes through. The fields map to nothing and the side effects never fired.

Your next prompt repeats the cycle, building on the model's own half-remembered inventions. Nothing enforces consistency between sessions, and no type system or schema check catches the errors. LLMs are unreliable narrators. An agent saying the work is done isn't evidence that it is.

Diagram showing how prompts drift over time without constraints

Rigid Layers

Think of your codebase as a stack of layers, some loose and some rigid. A loose layer is where the agent writes freely: application code, a schema, whatever the task calls for. A rigid layer is a rule that code has to obey—one that forces it to behave a certain way, or to hold a property you can check automatically. Loose layers are where the work gets made, and rigid layers are where it gets checked.

A type is the simplest rigid layer. If your schema says a payment carries an integer amount_cents, code that invents a total field won't compile, and the agent can't call the job done when the code doesn't typecheck. The silent failure gets caught before it ships. A Postgres CHECK (amount >= 0) works the same way. It rejects a bad row whether a human or an agent wrote the insert.

Rigidity is a spectrum. A type only checks shape. A test checks behavior, like whether the sender's balance comes out right after a transfer. A validator can go further still and encode a business rule directly, rejecting any transaction whose debits and credits don't balance. The more a layer can express, the more of the agent's work it actually verifies, instead of just confirming that the code runs.

The agent can write these types, tests, and validators itself. Each one it writes locks in intent that constrains whatever it writes next, including code from the same agent in a later session. But an agent that writes the check can also write a weak one, so the chain has to bottom out in a layer the agent can't rewrite, like the compiler or the platform you build on.

Diagram showing unconstrained AI output drifting away from what was intended

Neither kind of layer works on its own. A system made only of loose layers does the work but drifts out of correctness. A system made only of rigid layers is all constraints and no code.

Diagram showing alternating rigid and loose layers

That mix is what Deterministic Boundary Programming comes down to. The agent generates freely in the loose layers, and everything it produces has to pass through the rigid layers before it ships.


How

FRAGMENT is a ledger API for building financial products, and it stacks these layers for you.

You describe your money flows in a schema, listing the entry types, the accounts they touch, and how money moves between them. You or your agent write that schema, so it's a loose layer—but it answers to a rigid one beneath it. FRAGMENT checks the schema against the rules of double-entry accounting, and an entry type whose debits and credits don't balance is rejected before it exists.

From that schema, FRAGMENT generates a type-safe SDK, and the SDK becomes the rigid layer for the code above it. Your agent writes the business logic that posts entries, but it can only post the entries your schema defined, in the shapes it defined them. The agent writes both the schema and the business logic, yet it never writes the accounting rules beneath them—so however wrong it goes, the books still have to balance.

Tests sit at the top of the stack, and they're the one layer worth a tip. Write them adversarially. Have a fresh session that sees only your requirements, never the implementation, so the model isn't checking its own work.

Loose / Rigid
Testsrigid✓
App Codeloose✗
SDK / Typesrigid✓
Ledger Setuploose✗
Schemarigid✓
Schema Draftloose✗
FRAGMENT APIrigid✓

The quickstart walks through the schema-to-SDK loop, and Embed in CI shows how to make it a checkpoint your pipeline enforces by regenerating the SDK and failing the build if anything drifts.


The Rule

No AI output should become part of your system until it passes through a deterministic checkpoint. On any stack, those checkpoints are schemas, generated types, validators, and tests.