Inside the Loop, Part 2: The Guardrails That Let Me Trust an Unattended AI Agent
Inside the Loop, Part 2: The Guardrails That Let Me Trust an Unattended AI Agent
In Part 1, I watched my agent’s first real run line by line before I trusted it with a daily schedule. That answered “can it actually do the thing.” It didn’t answer the harder question: what happens on the days I’m not watching.
An agent that runs once, supervised, is a demo. An agent that runs every morning at 5AM, unattended, indefinitely, is a system — and systems need limits, not good intentions. This post is the actual configuration: the prompt, the permissions, and the specific lines I added after thinking through what could go wrong.
The Job, In One Sentence
Pick one unblocked, high-value ticket from my Linear backlog, do real work on it, open a pull request, then stop. Every morning. Forever, until I turn it off.
“Do real work” is doing a lot there, so let me break down what that actually means in the prompt, and then the limits wrapped around it.
The Cadence
5:00 AM Eastern, daily. Not because earlier is better — because I wanted it done before my own workday starts, so I walk into my morning with something already moved instead of a blank backlog staring at me.
Cloud routines run on cron in UTC, so this is 0 9 * * * (EDT is UTC-4 in August — I’ll need to remember to shift it when daylight saving ends). One agent, one repo, one model (claude-sonnet-5), a fresh sandboxed checkout every single run. No state carries over between days except whatever’s actually committed to the repo or written to Linear.
The Prompt: What It’s Told To Do
Stripped down, the daily instructions are:
- Read
CLAUDE.mdfirst. Company context, workflow conventions, confidentiality rules — all of it lives in the repo, not in the agent’s memory, so it’s current every single run. - List the backlog via Linear MCP, prioritizing the active product. Not everything I’ve ever created a ticket for — the live thing I’m actually building this quarter.
- Pick ONE item. Don’t try to clear the backlog. This is the rule I almost didn’t write, and it turned out to matter the most. More on that below.
- Skip anything already waiting on my reply, and anything requiring a live browser session with stored credentials (our community platform, Microsoft Teams). Those need a human at a keyboard. Don’t pretend otherwise.
- Move the ticket to In Progress and post a plan before touching anything. So if I check in mid-run, I know what it’s doing and why.
- Do the actual work — research, code, docs. Not a status update that looks like progress.
- On a genuine judgment call, don’t guess and don’t stall. Draft the best-reasoned default, flag it explicitly as a comment, and move on to something else instead of blocking the whole run on my answer.
- Push to a new branch and open a PR. Mark the ticket In Review, never Done.
- End with a plain-language summary: what moved, what needs me, what’s blocked.
None of that is exotic. What makes it safe to leave running is everything underneath it.
The Guardrails, and Why Each One Exists
”Pick ONE item, don’t clear the backlog”
The first draft of this prompt didn’t have this line. I added it after realizing that an agent optimizing for “make progress” with no ceiling will happily try to plow through ten tickets in one sitting, shipping ten shallow half-fixes instead of one solid one. Constraining scope per run was the single highest-leverage guardrail I wrote — it forces depth over breadth, and it means a bad day is contained to one ticket, not the whole backlog.
Why it matters: blast radius. If the agent misreads a ticket, the damage is one PR, not ten.
Never commit to main, never merge its own PR
Every change lands on a branch. Every branch becomes a PR. I review and merge, or I don’t. This isn’t a suggestion in the prompt — it’s a hard rule stated the same way for every agent that touches this repo, human-authored code included.
Why it matters: an agent’s confidence in its own output is not review. The whole point of the loop is throughput on my terms, not autonomy over what ships to production.
Cap of 3 open unreviewed PRs, more only if it’s small and low-risk
The agent isn’t allowed to bury me in a backlog of its own PRs while I sleep. Three open and waiting for me is the ceiling under normal circumstances. I loosened it slightly — “more if highly confident and the work is small, well-scoped, and low-risk” — because a rigid cap on genuinely trivial fixes (a typo, a broken link) would just create artificial friction. But the default is conservative.
Why it matters: review queue is a real cost. An agent that can generate infinite work faster than I can review it isn’t saving me time, it’s shifting the bottleneck onto my mornings in a worse form.
Skip anything already flagged and waiting on me — don’t re-nag
If a ticket has a prior comment asking me a question I haven’t answered yet, the agent checks for that and leaves it alone rather than pestering the same open question in a new comment every morning.
Why it matters: an agent that repeats itself daily trains me to ignore it. Silence on an unanswered question is more useful than noise.
Skip anything requiring live credentialed browser sessions
Our community platform and Teams meeting scheduling both require an interactive, authenticated browser session. The agent doesn’t have that, and I didn’t want it improvising a workaround. Those tickets get explicitly marked “blocked by tooling” instead of silently skipped or, worse, attempted badly.
Why it matters: the honest failure mode here is “I can’t do this, here’s why” — not a half-attempt that looks done but isn’t.
Judgment calls get flagged, not guessed or stalled on
This is the guardrail that did the most real work on day one. The agent hit an actual judgment call — a scheduled session time that didn’t match between the public landing page and the real calendar invite — and instead of picking one silently or freezing the whole run waiting for me, it applied a reasoned default (trust the side that’s already gone out to real people), made the change, and wrote out its reasoning as a comment so I could override it in five seconds if I disagreed.
Why it matters: the two failure modes I was actually worried about were an agent that silently makes calls I’d disagree with, and an agent that grinds to a halt every time it hits ambiguity. This threads between both — it acts, but it shows its work and leaves the door open.
Mark In Review, never Done
Only I close a ticket. The agent’s job ends at “here’s a PR, ready for your eyes” — it doesn’t get to declare victory on my behalf.
Why it matters: “Done” is a judgment about whether the work actually solves the problem, not just whether code got written. That’s mine to make.
What’s Deliberately Not Guardrailed
I didn’t restrict which ticket it can pick beyond priority and the two hard skip conditions above. I want it exercising real judgment about what’s actually high-value today, not just working top-to-bottom off a static list. If that turns out to be a mistake — if it keeps picking the wrong things — that’s a guardrail I’ll add later, based on evidence, not one I’m guessing I need now.
That’s the pattern across all of this, honestly: every rule above exists because I thought through a specific failure mode, not because I imagined every possible one in advance. I’d rather ship a tighter loop with fewer, well-reasoned limits than a paranoid one with rules for scenarios that never happen.
What You Can Do Today
- Write your “pick one” constraint before your first run, not after. Unscoped agents default to doing as much as they can, and more isn’t always better.
- Decide your merge boundary in advance. Mine is: the agent opens, I merge. Non-negotiable, stated identically everywhere the agent might read instructions.
- Give it a way to flag a judgment call instead of only “ask” or “guess.” A default plus a visible flag beats both a silent decision and a stalled run.
Part 3 is the actual results — how many PRs, how many judgment calls, what I had to correct. That needs a few real mornings behind it first, so it’s coming once there’s real data to report, not before.
Comments