Harness Engineering: Turning an AI Agent into a Reliable Teammate
A coding agent is only as good as the environment you put it in. The model brings raw capability; the harness — the context, tools, guardrails and workflow around it — is what turns that capability into code you can actually merge into a production system.
For the past months I've been running an AI coding agent (Claude Code) daily on a multi-tenant SaaS backend built with Laravel 11, Octane/Swoole and PostgreSQL. The codebase has strict rules: tenant isolation, a fixed API response contract with an external front-end, and stateless code that survives long-lived workers. Early on, the agent produced code that looked right and quietly broke those rules. The fix wasn't a better prompt. It was a better harness. This is how I built it, layer by layer.
Table of contents:
What Is a Harness?
The harness is everything that surrounds the model: what it reads before it starts, which tools it can call, which checks run whether it likes it or not, and which steps it must follow from task to merge. Two engineers using the same model can get completely different results, and the difference almost always lives here.
The guiding principle I follow is simple: anything that must always happen should not depend on the model remembering to do it. Instructions guide behavior; scripts enforce it. The harness is the place where you decide which is which.
Layer 1: A Lean Project Memory
The first file the agent loads is `CLAUDE.md`. The temptation is to dump the whole architecture there. That backfires: a long memory file dilutes the rules that matter and costs context on every single turn. Mine is about thirty lines and does three things only:
- Points to the source of truth: full conventions live in `ARCHITECTURE.md` and domain docs in `docs/<resource>.md`. The agent is told to read the relevant section, not everything.
- Lists the exact commands, including the traps. Tests run inside the container with `docker exec`, and inside a worktree they need `-w <path>` — otherwise the suite runs against the main checkout and passes without ever seeing your changes. That one line has saved me from more false greens than any test.
- States the non-negotiable rules: tenancy through a global scope instead of manual `where('tenant_id')`, a single response builder instead of raw JSON, validation only in Form Requests, queries only in repositories, and no request state in `static` or singletons under Octane.
If a rule needs a paragraph of reasoning, the reasoning goes to a doc and the memory keeps a one-line pointer.
Layer 2: Skills as Reusable Playbooks
Skills are packaged instructions the agent loads only when a task calls for them. They keep the base context small while making specialized knowledge available on demand. The ones that earned their place in this project:
- Conventions: the execution and review checklist for any PHP change — the exact order to build an endpoint (route, Form Request, DTO, controller, service, repository, resource, binding, Pest test) and the rules that are never negotiated.
- Spec & Plan: any change with more than three steps produces a written spec and plan in `specs/plans/` before a single line of code. The agent explores, specifies, asks for approval, then executes step by step.
- Create Task: turns a raw request — an audio, a screenshot, a client message — into a product-level issue: current flow, desired flow, business rules and acceptance criteria. Technical suggestions are forbidden by design, so the planning step stays free to choose the right solution.
- Post-Task Finalizer & Doc Generator: after the work is done, they draft the issue comment for the front-end team and update the internal docs, based only on facts from the real diff.
- Commit: Conventional Commits in English, and an absolute rule that commit messages belong to the developer — no AI credits or signatures.
Layer 3: Hooks as Deterministic Guardrails
Skills and memory are still requests. Hooks are not. A hook is a script the harness runs at a fixed point of the agent's lifecycle, and its exit code can block the agent from moving on.
The Stop Hook: Anti-Pattern Scan
Every time the agent tries to finish, a Bash script scans the diff of `app/` and `routes/` and exits with code `2` if the added lines introduce any of the patterns the architecture forbids:
- Manual `where('tenant_id', ...)` outside deliberate cross-tenant flows.
- Raw `response()->json(...)` instead of the shared response builder.
- A global `Cache::flush()` in a multi-tenant system.
- `auth()->user()` inside services, repositories, models or jobs, where a request-scoped tenant context must be injected.
Design Decisions That Made It Work
The scan only looks at added lines, never at existing code. Legacy that still uses an old pattern doesn't fire; the hook only prevents adding more of it. Exceptions are an explicit, short allow-list in the script itself — each one documented with the reason and the plan that justified it, because an exception is debt, not permission. And the message tells the agent exactly how to fix the problem, so it fixes it instead of working around it.
Layer 4: Commands and Sub-Agents
The top layer ties everything into a workflow. A single slash command, `/resolve-issue <number>`, drives an issue from the board to a reviewed commit:
- Isolate the work in a dedicated git worktree and branch.
- Read the issue through a sub-agent that operates the GitHub Projects board, summarize it, and move the card to In Progress — which doubles as a lock so another session doesn't pick up the same task.
- Write the spec and plan when the task is non-trivial.
- Implement with tests for every new endpoint, branch or rule; bug fixes require a regression test. The whole suite must be green.
- Re-check the full diff against the Octane-safety and convention rules, line by line.
- Run an independent review: a fresh sub-agent that doesn't inherit the author's reasoning reads the spec, the diff and the conventions, and reports only gaps in correctness, security, tenant isolation or the API contract.
- Draft the docs and the issue comment, then stop and show me everything — including a self-review checklist with evidence — before anything is published or committed.
The board operator is its own sub-agent with the project's field and option IDs already mapped, so moving cards, assigning sprints and commenting don't pollute the main context with GraphQL exploration.
Running Agents in Parallel with Worktrees
Once one session works reliably, the next step is running several at once. Git worktrees give each issue its own directory and branch, so edits, `git status` and the stop hook stay scoped to that task. Getting there taught me a few non-obvious lessons:
Copy dependencies, don't symlink them. Both Composer's autoloader and Node's module resolution compute the project base from the physical path. A symlinked `vendor/` silently resolves back to the main repository, so the worktree ends up running main's code — and tests pass or fail for the wrong reason. A copy-on-write clone (`cp -c` on APFS) makes a real copy almost instantly without duplicating disk space.
Share the harness, not the state. The `.claude/` folder is symlinked into every worktree, so all sessions follow the same rules, while `.env` is copied as a plain file because a host path would not exist inside the container.
Know your shared resources. The development database is shared across worktrees, so full test runs from parallel sessions can collide. The rule is written down: when running in parallel, run suites serially.
Final Thoughts
Working with AI agents in production is less about prompting and more about engineering the system around the model. Every time the agent made a mistake, I asked one question: is this something to explain, or something to enforce? Explanations went into memory and skills; enforcement went into hooks and workflow steps. Over time, the same model went from producing plausible code to producing code that follows the architecture by default.
- Keep memory lean: Short rules and pointers beat long documents the model has to sift through on every turn.
- Enforce what matters: If breaking a rule causes a real incident, a script should block it — not a sentence in a prompt.
- Separate author and reviewer: A fresh agent with no inherited reasoning catches what the author rationalized away.
- Keep a human gate: Nothing is published or committed without explicit approval. The agent does the work; the engineer owns the result.
Let's talk about your project!