Every guard is a rule in the wrong place

I spent a Tuesday writing a guard so that one guard would stop contradicting another. What ended the approach was a second model auditing the first: it fired eighteen times at seventeen per cent precision, and the week after it moved every rule out of the prompt and into code, where a test can reach them.

I spent a Tuesday writing a guard so that one guard would stop contradicting another guard.

What they were guarding is a chat agent that runs a structured intake conversation. It asks about your trade, your languages, where you will travel, whether you need housing, eleven fields in all, then proposes what fits from a live catalogue. Nine tools behind it, and one conversation that has to feel like a person.

About twenty-five deterministic backstops sat downstream of the model, correcting its prose before any candidate saw it. Eleven were text-mutating passes with names like stripLeakedToolCalls and enforceOneQuestion, and two of them had begun disagreeing over whether a sentence had already been rewritten.

Here is the claim: a guard is a rule living somewhere no test can reach, and a system that keeps growing guards is telling you where its policy sits rather than how careful its author is. Every one of mine ran after the decision, on prose, in a pass that could only see the finished string. Rules at that layer need rules about the rules, which is a thing I had to build before I believed it.

A backstop is a rule that runs after the decision

For four months I thought an agent was a castrated ChatGPT. Pin a general model to a role, hand it a fixed toolset, and you have an employee who never sleeps. The first version was exactly that, and it worked beautifully right up until real people started answering it.

[past-max]

Real people give a birth year when you ask their age. They send two phone numbers, or one with a digit missing. Some have names that reveal no gender, so the bot has to ask without making it strange. Some ask a question without a question mark, and the code downstream has to guess whether to answer or move on.

Each of those became a line in the system prompt, until the prompt grew long enough that the model began ignoring parts of it. So the architecture became a slogan: the AI reacts, the code checks. Before anything reached a candidate, code inspected the draft, recorded what the model had failed to record, and nudged it into phrasing the next question.

For a few weeks this felt like power. Any whim could become another guard.

Then the deviations kept arriving. Someone gives a phone number only at the end, once a vacancy looks worth it. Someone changes country mid-conversation. Someone is job-hunting for themselves and their husband, in one thread. A fragment of a tool call leaks into the prose. The bot addresses a candidate whose name it has yet to learn, using its own.

Every fix landed in the same place, because that was the only place left: after the model had written the sentence, on the sentence.

The strongest case for guards is that every one of them was real

Each backstop had a live incident behind it and a comment naming the incident. None was speculative, and taken one at a time each was the cheapest correct fix available that day. Anyone arguing that I should have designed better up front is arguing against how the requirements actually arrived, which was one candidate at a time.

The auditor was the best idea in the whole approach. On 14 July I gave a cheap second model the draft reply plus the ground truth of what had really happened that turn, the actual tool calls, and asked whether the text claimed an action that never ran. Semantic, so paraphrase could not slip past it. Language-agnostic, so it would survive the agent answering in any of the languages my patterns were never written for. On its first live evaluation it caught a sentence along the lines of “let me check a few more details”, a promise of future work in a phrasing none of my patterns had ever seen.

That is the version of the objection I want to lose to, and for about six hours I was winning.

Precision is the thing a corrector has to earn

Measured over a fifteen-persona run, the auditor fired eighteen times at roughly seventeen per cent precision. Three of the sentences it “repaired”:

My first hypothesis was missing context, so I gave the auditor tool outcomes instead of tool names. Logging what it actually received settled that: the outcomes were there and it flagged anyway. A small model asked to hold that many carve-outs will drop some of them, and a false correction costs more than a missed one, because it damages the natural voice the whole design exists to protect.

[devil's-advocate]

So I narrowed it to exactly one question: does the bot promise to do something itself, later? That is also the only thing the deterministic backstops could not already do. Measured with npm run audit:probe, the narrow version scored eight out of eight on both recall and precision, including two self-promises made in languages my patterns had never been written to cover. Fires across the evaluation went from eighteen to zero.

Then I looked at what I had. A model whose entire job was to distrust a model, wrapped in guards of its own, asking one question that a typed field would answer for free.

Moving the decision deletes the failure instead of catching it

Two days later the rewrite landed: a net loss of fifteen hundred lines, and three changes.

The model no longer writes the final message. It returns parts, which are what it meant, how it reacted, the answer to the candidate’s question if there was one, and how the next form question should be put. Code assembles the string. Leaked tool calls, two questions in one breath, a sentence cut off halfway: there is no longer a step in the turn where any of those could be produced, so the passes that used to remove them have nothing to do.

The model no longer decides which field is missing. A pure function does, so nobody is asked the same thing twice.

Every claim carries an intent label from a closed set of eleven, and code folds the turn’s tool trace and compares. A reply announcing that it found some roles with no search in the trace either triggers the search or gets downgraded to something honest. The remedy keeps the model’s answer verbatim and swaps only the false claim, which is what the prose surgery had been attempting all along, now done by construction.

where the policy lives same agent, two architectures
the model drafts, code corrects
  • every rule lives in the system prompt, which the model partly ignores once it is long enough
  • the model writes the whole final message
  • the model decides which field is still missing
  • ~25 prose backstops rewrite the draft before it ships
  • a second model audits the first, and needs guards of its own
code decides, the model phrases
  • a cheap NLU pass reads the message before any generation runs
  • the model returns parts; code assembles the string
  • a pure function picks the next field
  • a closed set of 11 intents, each checked against the turn's tool trace
  • a false claim is a typed field, so the auditor has nothing left to read
one turn, after the rewrite the model appears twice, and never last
  1. ● 01
    reconcile
    cheap NLU pass, before any generation
  2. ● 02
    plan
    planIntakeMove(): code picks the field
  3. ● 03
    generate
    flat reply: intent, reaction, answer, question
  4. ● 04
    verify
    intent checked against the tool trace
  5. ● 05
    compose
    composeReply(): the only place the string is built

Changing all this was frightening, because the old thing did work. So, tests: 804 green the evening the structured loop landed, 1,190 by the time I wrote this. Ninety-seven simulated candidates, sixteen turns each, another model playing the human deliberately badly, answering only what was asked, one fact at a time, the way people do and fixtures never do. One spends four turns trying to break the agent, fails, then settles down and looks for a job like everybody else. One rule in that harness exists purely so a simulated candidate will sometimes skip the introduction and open with a question, which is the case that used to leave the agent addressing a stranger by the only name in the room. Its own.

[max]

What moving policy into code does not buy

The system prompt came out of this longer than it went in. I had assumed that moving policy into code would shrink it. What moved was the load-bearing part, into somewhere a test can reach, which is a different win from the one I went looking for.

The claim also needs a policy to exist before it helps. This agent talks to a candidate and serves the company: it gives feedback in natural language and moves a script along. Solving anybody’s problem cleverly was outside the brief the entire time, which is why the rules could be enumerated at all.

[devil's-advocate]

Yes. An agent whose job is genuinely open, where the user’s goal is the thing being served, has no equivalent list to move into code, and everything here would have to be argued again from the start.

What I have is convincing hope that this one will behave predictably and stay cheap to extend. Sooner or later we will overload it again with the variability and mutual contradiction of our own requirements, which we track badly, and it will start growing guards again. The difference is that I now read a new guard as a question about where its rule belongs, and I have yet to work out how many of them have to appear before I am obliged to ask it.