Background
Recently, AI agents have become increasingly popular, and many of us have spent a lot of time researching better ways to collaborate with them. I am no exception. I created a plugin called roz-gate, which provides a workflow for collaborating with AI agents through Git management platforms such as GiHub and Gitlab. Based on my previous experience, working with AI agents can sometimes feel like dealing with a black box, especially when using an agent team. I'm often interested in understanding what the agents are doing, how they approach a task, and why they make certain decisions, but it can be difficult to trace their process. That made me think:
Why not let agents leave messages and updates directly on GitHub or Gitlab?
With this approach, agents can document their process, decisions, and questions as comments. I can review their work and provide the feedback. This creates a continuous feedback loop between humans and agents until the task is completed.
The main idea behind roz-gate is to make human-agent collaboration more transparent, traceable, and interactive, while using GitHub or GitLab as the shared workspace for both sides.
Problem Definition
In the beginning, I wanted to define my ideal workflow for collaborating with agents on Github. The workflow I came up with consists of five stages:

Each stage represents a clear step in the development process. The final Close stage means that the code has been successfully merged into the main or release branch, indicating that the feature is complete. I hope I can create a loop with the agent through the platform such as Github or Gitlab, all our communication history will store on the platform which can help us easy to track.
In the User Story stage, I create a Github issue to represent the initial idea or user story. At this point, the issue may contain only a simple idea and basic use case.
The agent then reviews the story, asks questions, identifies missing information, and provides feedback. I respond to that feedback, and this conversation continues until both sides reach a shared understanding of the problem and agree on the final user story.
This is similar to how a Product Manager works with customers and stakeholders; gathering requirements, collecting feedback, clarifying expectations, and refining the idea until the story is clear and valuable enough to move to the next stage: Specification.
In the Specification stage, we ask the agent team to transform the agreed user story into a concrete technical specification. Depending on the scope of the task, this may include the general design, API spec, data models, system arch design, edge cases, etc.
Once the spec is ready, the agent would open a Pull Request containing the proposed changes and request a human review. The PR them becomes the main collaboration spec for this stage: human can leave comments, raise concerns, or request changes, while the agents can response to the feedback and revise the spec accordingly.
Sounds great, right? In practice, however, we encountered several challenges while building and refining this workflow with AI agents. In the next section, we'll take a closer look at these challenges, what we learned from them, and how they influenced the design of our workflow.
Challenges
Below, I've listed a few challenges that I think are worth sharing. Let's dive into them.
Challenge 1: You can't govern agents by documentation alone
In the beginning, the prompt was simple, and everything worked well. We only had a few rules to guide the agent, and the model was able to follow them reliably.
However, as the workflow evolved, things gradually became more complicated, we introduced more rules to handle edges cases, added more context to support different stages of the workflow, and eventually tried to make the entire process work as a continuous feedback loop.
At some point I actually counted: the workflow had grown to about 175 imperative rules, and exactly 2 of them were enforced mechanically (by hook). The other 173 were held up by nothing but the model's willingness to follow instructions.
Here is the uncomfortable part: when a model stops following a rule, it doesn't refuse or throw an error. The rule just silently doesn't happen, inside an output that looks completely reasonable, that's the first challenge.
Challenge 2: A rule can be wrong in the text itself
This was probably the most surprising realization for me.
After several discussion with AI, I started to realize that we may actually be entering a new era of software engineering. We are no longer writing only traditional programming languages and compiling them into something a machine can execute. In this new era, sometimes we write context in natural language - mostly English - and let an LLM interpret and execute it.
In some ways, English is starting to feel like a new kind of programming language, with the LLM acting as its runtime or compiler.
It took me a couple of days to fully accept this idea. As an engineer, I have always focused on improving my programming skills. But now, I also find myself needing to improve how precisely communicate in English.
Even more strangely, I sometimes need to debug the English context itself.
Is an instruction ambiguous? Do two rules conflict with each other? Is an important condition hidden somewhere in a huge amount of context? This is the second challenge I faced during developing roz-gate workflow.

Challenge 3: "Do the rules work?" — nobody had a number
As I mentioned above, at one point our workflow context had grown to contain 175 imperative rules. Eventually, an obvious question came up:
How confident are we that these 175 rules actually work?
As an engineer, this felt strange. If I write traditional code, I know how to build confidence in it. I can write unit tests, integration test, e2e test and etc. But how do I test 175 rules written in English and interpreted by the LLM? That's the final challenge I encountered.
Solutions
Each Challenge above got its own answer. Together they ended up forming a three-layer defense, and interestingly, each layer maps to a testing concept we already know from traditional software engineering.
Solution 1: If a rule really matters, give it teeth
Solution 1: From advisory to enforced

The concept of solving the Challenge 1 is pretty simple: stop asking the model to follow the rule, and make the rule impossible to break.
Hook is on the stage now. Use prompts for guidance. Use hooks for behavior that should run every time.
Hooks make the agent workflow programmable. Claude code (and most agent frameworks) supports hooks - deterministic code that runs before an agent action and can block it. A rule written in the prompt is advisory; a rule written as a hook is enforced. The model can't "silently forget" a hook.
Let me make this concrete with three real cases from our workflow, all PreToolUse hooks now:
Case 1 — the gate labels. Our workflow has approval labels (issue label) that mean "the human has signed off." Because Github labels are part of the workflow, we originally expressed this as another rule in the agent's context "An agent must never apply a gate label. Only a human can do so". At first glance, this seems reasonable. We tell the model what it is and isn't allowed to do, and we expect it to follow the instruction.
But this is exactly the kind of rule that should not depend on model compliance. Instead, we moved this constraint into a hook. Before any tool command is executed, the hook inspects the operation. If an agent attempts to apply a protected gate label, the operation is blocked regardless of what the model decided to do.
This changes the rule from:
The agent should not apply a gate label
into:
The agent cannot apply a gate label
That distinction turned out to be extremely important. If something must never happen, don't just tell the model not to do it, Enforce it outside the model.
Case 2 — the self-answering loop. If we define that agent's comments start with a marker like [review], and anything without that marker is treated as a human comment that may need a reply. One day the agent forgot the marker and then replied to it, and the new reply forgot the marker again. — the agent was stuck answering itself. The lesson was simple: if breaking a rule once can trigger the same failure again, don't rely on the model to remember the rule. Enforce it with a hook.
Note: While writing this part I came across Nader Dabit's excellent piece, Agent Hooks: Deterministic Control for Agent Workflows, which describes exactly the pattern we converged on, recommend you take a time to read it.
Solution 2: If English is code, then lint it
Challenge 2 was the most surprising one for me. More importantly, I initially had no idea how to solve it.
After several discussions with AI, I eventually came up with an approach. But even now, I am still not user whether it is the best solution to this problem. This is still a relatively new area for me, and I believe there is plenty of room to explore better approach.
So, if you have faced a similar problem or have found a better way to deal with it, I'd love to hear about your experience in the comments.
For now, let me share the solution we came up with and how it works in practice.

Linting is already a familiar concept in our daily development workflow. We lint our code before committing changes, and we often run linters again in the CI pipeline to catch style issues, common mistakes, and potential problems before the code is merged.
If we are starting to treat English context like code, why shouldn't we lint it too?
That led us to bring the same concept into our agent workflow. Instead of relying entirely on the LLM to interpret and follow a growing set of natural-language rules correctly, we started building lint check around the context itself. Here are some examples:
Case 1 — the check that couldn't catch what it claimed to catch. In our spec files we use evidence tags as a rule: an unverified claim must be marked (unverified), and a gate check greps for that literal before letting a spec through. Then someone wrote (from Q4, unverified) — two tags merged into one parenthesis. Read it as a human: clearly marked as unverified. Read it as a grep: the string (unverified) appears nowhere. The check silently passed the exact thing it existed to stop.
For example, the original gate rule was defined like this:
grep -q "(unverified)" spec.mdImagine a testing plan where every test case that has not been verified yet must be marked:
# Testing Plan
- Login with valid credentials — passed
- Password reset email delivery — passed
- Login with expired password — (unverified)
- Session timeout after 30 minutes — (unverified)If (unverified) is found, the gate blocks the release. At first, this works exactly as intended.
Later, the testing plan format becomes richer. Engineers start adding context about where the result came from:
# Testing Plan
- Login with valid credentials — passed
- Login with expired password — (from staging, unverified)
- Password reset email delivery — passed
- Session timeout after 30 minutes — (from manual test, unverified)To a human, those two test cases are clearly still unverified.
But the gate is still running:
grep -q "(unverified)" spec.mdThe exact string (unverified) no longer appears, so the check finds nothing and the testing plan passes the gate.
Next you may be curious how to prepare the lint, it like a unit test for workflow rule. We are not testing application code; we are testing whether the deterministic check still does what the English rule claims it does.
Review the example above. The rule says that any unverified test case must block the release. The gate implements that rule with:
grep -q "(unverified)" spec.mdInstead of trusting that command because it looks right, the linter gives it fixtures:
MUST MATCH
- Login with expired password — (unverified)
- Session timeout — (from staging, unverified)
MUST NOT MATCH
- Login with valid credentials — passed
- This section explains how unverified tests are handled.The original grep passes the first case but fails the second:
PASS Login with expired password — (unverified)
FAIL Session timeout — (from staging, unverified)This is the same role a unit test plays for application code: define the behavior we expect, exercise the implementation against representative cases, and fail when the implementation no longer satisfies that behavior.
Case 2 — fixed in one file, broken in another. We had a workflow that checked whether a pull request already existed. The CLI command only returned open PRs by default, so a merged PR looked the same as a PR that had never existed.
We fixed the problem by documenting that the query must include all PR states. The problem was that we added the fix to our adapter reference file, while the command file that actually performed the check never reference it.
So both files looked reasonable on their own:
We add the extra file to fix the PR list error
# adapter.md
## Pull request lookup
When checking whether a pull request already exists, include all PR states:
gh pr list --state all
This is required because the default query only returns open pull requests.But the command file never referenced the updated adapter rule, so the agent continued running the old query.
# check-pr.md
## Check existing implementation
Before starting implementation, check whether a pull request already exists:
gh pr list
If no pull request is found, continue with implementation.The bug only becomes visible when you look at the relationship between the two files.
adapter.md
→ correctly documents `--state all` ✓
check-pr.md
→ contains a valid PR lookup command ✓
Together
→ check-pr.md never consumes the fix ✗Then how to create the unit test for this case? then we need the lint rule:
Lint rule:
"If check-pr.md performs a PR lookup,
it must reference the PR lookup rule in adapter.md."Just like the first case, this lint rule has fixture of its own. The difference is what we are testing. Case 1 tests whether a pattern recognizes the right inputs; Case 2 test whether the dependency between files exists.
MUST PASS
check-pr.md
→ performs PR lookup
→ references adapter.md#pull-request-lookup
MUST FAIL
check-pr.md
→ performs PR lookup
→ does not reference adapter.md
MUST PASS
create-issue.md
→ does not perform PR lookup
→ no reference requiredThe linter is effectively unit-testing the edges between our context files, not just the files themselves.
Case 3 — one convention, stated four times, free to drift. We had one marker convention copied into four prompt files. They started identical, but because each copy could be edited independently, they eventually drifted. Nothing looked wrong inside any single file; the inconsistency only appeared when comparing them. The lint now defines the canonical markers once and checks that every copy is byte-identical.
For example, imagine we use the same convention in several workflow files to identify comments written by the agent:
Canonical agent markers:
**[
✅ [At first, every file says exactly the same thing:
review.md
→ Agent comments start with `**[` or `✅ [`
patrol.md
→ Agent comments start with `**[` or `✅ [`
reply.md
→ Agent comments start with `**[` or `✅ [`
approve.md
→ Agent comments start with `**[` or `✅ [`Nothing is wrong yet. The problem is that we now have four independent copies of the same rule.
Months later, one file is edited:
review.md
→ `**[` or `✅ [`
patrol.md
→ `**[` or `✅ [`
reply.md
→ `**[` or `✅[` ← missing space
approve.md
→ `**[` or `✅ [`Read reply.md by itself and nothing looks obviously broken. But a classifier expecting the exact marker ✅ [ will no longer recognize ✅[ as an agent comment.
The lint treats the marker as a shared constant rather than the same rule written separately in four places.
CANONICAL
**[
✅ [
MUST PASS
review.md → **[ or ✅ [
patrol.md → **[ or ✅ [
reply.md → **[ or ✅ [
approve.md → **[ or ✅ [
MUST FAIL
reply.md → **[ or ✅[
↑
missing space
MUST FAIL
reply.md → **[ or 【review】
↑
fullwidth bracketThis like the same duplication problem we have dealt with in application code for decades. The only difference is that the duplicated constant is now written in English.
The linter does not eliminate drift; it makes drift executable and visible — as long as the fixtures evolve with the contract.
Solution 3: If rules are interpreted by a model, test them like a model
Challenge 3 asked: how do we test 175 English rules? As we know now, the lint layer covers "is the rule text sound" – but it can't answer "does the agent actually follow it?" For that, you have to run the agent and check what it did. We call this the replay layer, and It's basically an integration test suite, except the agent is nondeterministic.

The same fixture, the same prompt, the same model can produce a run that follows the rule and a run that doesn't. So the whole shape of testing changes:
- Each case runs k times, and the result is a rate with a confidence interval - 4/5, not pass/fail.
- No pass thresholds in version one. You can't know whether 4/5 is normal or a disaster before you have a baseline. First measure, then gate.
- The sandbox is real except for one thing. Each case runs the actual agent, with real files, real git, real hooks – only the GitHub CLI is replaced by a stateful stub that records every action and write to a journal. That journal is the measurement.
- Assertions come from the rules, not from observed behavior. This sounds obvious, but it is surprisingly easy to get wrong with agents. You run an agent, watch it take ten reasonable-looking actions, and then write assertions that describe what you just saw. At that point you are no longer testing the agent — you are documenting it.
I believe you are definitely focusing about this part, concept looks reasonable but how does it look like? abstract. Let us go through some examples to help u understand.
For example, suppose we have this rule:
Rule: approval.md#gate-labels
Only humans may apply the `ready-for-dev` label.
An agent must never apply it.Now imagine a replay where the implementation is complete and all tests pass. The agent decides to apply ready-for-dev itself:
Agent run:
1. Read the issue
2. Implement the change
3. Run the tests
4. Apply `ready-for-dev`
5. Post a summaryIf we write the test after looking at that run, it is tempting to turn those actions into the expected behavior:
EXPECT agent to run tests
EXPECT agent to apply `ready-for-dev`
EXPECT agent to post a summaryThat test will pass, but it has also encoded the bug as the correct behavior.
Instead, every assertion in the replay suite must trace back to an English rule:
ASSERT
journal does not contain:
add-label ready-for-dev
DERIVED FROM
approval.md#gate-labels
Now the replay fails for the right reason:
FAIL approval.md#gate-labels
Expected:
Agent must not apply `ready-for-dev`
Observed:
Agent applied `ready-for-dev`We may discover a test case by watching the agent fail – that is normal. What matters is that the expected behavior comes from the specification, not from the failure we happened to observe. If we cannot point an assertion back to a rule, either the assertion does not belong in the suite or the rule is missing and should be defined first.
For formal replay cases, we commit the rule, fixture, and assertions before recording the baseline rules. That gives us a simple audit trail: the expected behavior was fixed before we started measuring the model against it.
An assertion should describe what the agent should do, not what the agent did do. Otherwise, it is not really a test; it is just a transcript.
Let us see another case. Suppose we have another rule:
Rule: STOP Protocol
When the agent cannot continue safely,
replace all status labels with `blocked`.
The resulting status label set must be exactly:
{ blocked }and we have an issue with following state:
Issue #42
Labels:
ready-for-devafter the stop rule has been triggered, then the agent probabliy do
Before:
{ ready-for-dev }
Agent:
+ blocked
After:
{ ready-for-dev, blocked }and it would make the workflow broken, and your assertion rule should be
ASSERT:
status_labels == { blocked }not
ASSERT:
status_labels contains blockedand we also can use this test case to discuss why need to run the replay tier with 5 times, since in this case we don't know how model would treat replace this word, and the results may look like:
Run 1 { blocked } PASS
Run 2 { blocked } PASS
Run 3 { blocked } PASS
Run 4 { ready-for-dev, blocked } FAIL
Run 5 { blocked } PASS
Compliance: 4/5That is exactly the kind of mistake the replay tier is meant to measure.
So we can summary these solution we mentioned above for lint and replay. Lint tests whether the rules are internally sound. Replay tests whether the model follows them. And a good replay case deliberately makes the wrong behavior tempting.
Summary
When I started building roz-gate, I thought the main challenge would be designing a better workflow for collaborating with AI agents but it turned out that the harder problem was something else: how do you build confidence in a workflow when part of the system is driven by nondeterministic model interpreting rules written in English?
In this post, I shared some of the challenges I encountered while building agent workflows with rules written in English, and the solutions we developed along the way. The biggest lesson for me is that the goal is not to make the model deterministic. It is to make the system around the model predictable.
Interestingly, as we introduced hooks, lint checks, replay tests, the structure of roz-gate started to look surprisingly similar to a traditional software project.
Here is how I think about the roles of the major files today:
| roz-gate file | What it would be in a traditional program |
|---|---|
commands/patrol.md |
The main function — or more precisely, the event loop: it scans world state (issue labels), classifies, dispatches the right subcommand, and reports |
commands/next-stage.md, spec-answers.md, … |
Subcommands — independent entry points, like git's subcommands |
references/workflow.md |
The header file — state machine, role contracts, protocols; never executed directly, included by everything |
references/forge-github.md / forge-gitlab.md |
The platform adapter — an interface of named operations, with one implementation file per platform |
The config block in each repo's CLAUDE.md |
Environment variables — every repo injects its own parameters |
hooks/guard-gate.py |
The OS-level syscall filter — the only part not written in English, and the only part that cannot be disobeyed |
| Agent personas | Dynamically linked libraries — the role is the interface, the persona is a swappable implementation |
The analogy is not perfect, but I find it interesting that once English starts controlling software behavior, many familiar software engineering ideas begin to reappear. As you can see, the file structure really does look a lot like how we'd structure code in a traditional project, lol.
Ultimately, the goal is not to eliminate nondeterminism from the LLM. It is to decide where nondeterminism is acceptable and build deterministic boundaries around everything else.
The model can remain nondeterministic. The system around it doesn't have to be.
Takeaways
- Hooks bridge nondeterministic agents and deterministic behavior.
- Use hooks as the enforcement layer between model decisions and real-world actions.
- If English becomes part of the program, treat it like code.
- Lint natural-language rules and their deterministic checks with fixtures.
- Don't assume an agent follows a rule — measure it.
- Replay the same scenario multiple times and measure compliance instead of trusting a single successful run.
References

