Loop Engineering: Why I Watched My AI Agent's First Autonomous Run Line by Line

• by Alien Brain Trust • AI Learning
Loop Engineering: Why I Watched My AI Agent's First Autonomous Run Line by Line

Loop Engineering: Why I Watched My AI Agent’s First Autonomous Run Line by Line

I set up a cloud agent to work my company’s backlog every morning at 5AM. Pick one ticket, do real work, open a pull request, stop. By the time I’m at my desk, something’s already moved.

That’s the pitch. Here’s the part nobody puts in the pitch: the first time you hand an agent a recurring schedule and walk away, you’re trusting it with access you haven’t actually watched it use yet.

I almost shipped this on faith. I’m glad I didn’t.


The Question I Almost Skipped

A few days earlier, I’d connected Linear to my agent’s toolset. It didn’t work automatically — I had to go authorize it explicitly, as its own connector, with its own permissions, before any Linear tool call would succeed. That was fresh in my head when I sat down to wire up GitHub.

So I asked myself the obvious follow-up: does GitHub need the same treatment? Or does it just work because the agent already has repo access?

The honest answer was: I didn’t know. And “I didn’t know” is exactly the condition under which people ship things that fail silently. A daily routine that can’t actually push code isn’t a broken feature you notice — it’s a green checkmark that quietly does nothing, every morning, until you happen to look.

I could have guessed. Plenty of people would have. Instead I triggered the routine manually and read the log.


What “Just Watch It” Actually Looks Like

Not a summary. Not the agent telling me it worked. The raw event stream — every tool call, every result, in order, as it happened.

The first few lines were unremarkable in the best way:

Cloning repository base-bit/alienbraintrust-private
Finished processing sources
git status → nothing to commit, working tree clean
git branch -a → * main, remotes/origin/main

Good. Read access to the repo was live, and it didn’t come from a connector I’d set up — it came from the routine’s own repo configuration, the same mechanism that clones the code in the first place. One question answered: cloning and reading, solved by default, no extra plumbing.

Then it got interesting. The agent picked up a real ticket — one asking it to confirm that a scheduled session time was correct — and instead of taking the ticket’s word for it, it went and checked the actual calendar invites. That’s when it found something nobody had caught: three landing pages advertised Wednesday sessions at 8:00 PM. The real Google Calendar invites, already sent with live meeting links, said 9:00 PM. A full hour off, live on the public site, for months.

The agent didn’t ask me what to do. It fixed the two pages that hadn’t already run their affected session, left the one that had already happened alone, committed on a branch, and pushed.

That push is where my actual question got answered. The log showed a git push succeed over plain git — no MCP involved — and then, seconds later, a call to mcp__github__create_pull_request. A GitHub connector I had never manually configured had already attached itself to the routine, the same quiet way Linear’s had once I’d authorized it at the account level. It opened PR #96, checked for CI, found none configured for that path, and moved on.

Then one more thing happened that I hadn’t engineered at all: GitHub fired a webhook back into the same session the moment the PR went up, subscribing it to future activity — new comments, CI results — so it wouldn’t have to sit there watching. It responded by scheduling itself a check-in an hour later and closing out the run.

I didn’t build the webhook loop. It came with the connector. I just had to watch closely enough to see it happen.


What “Loop Engineering” Means to Me

I’ve spent 25 years in enterprise security doing variations of the same exercise: don’t assume a system has the access it needs, or the access it’s supposed to be limited to — verify it, on the actual system, under real conditions. Config review is a starting hypothesis. A live trace is the answer.

Autonomous agent loops deserve the identical treatment, and for the same reason: the failure mode isn’t a crash you notice. It’s quiet, plausible-looking non-work — or worse, quiet overreach — that you don’t catch until it’s cost you something.

Three things I now do before I trust any scheduled agent loop with real access:

1. Trigger it manually before the schedule ever fires. A cron expression tells you when something will run, not whether it can. The first real test should be one you’re watching, not one that happens while you’re asleep — which was the entire point of this routine.

2. Read the tool-call log, not the agent’s summary of itself. The summary at the end of a run is the agent’s account of what it did. The event log is the record. I wanted to see the actual git push and the actual PR-creation call, not a sentence asserting they happened. This is the same instinct as reading raw audit logs instead of trusting a status dashboard — the dashboard is a claim, the log is evidence.

3. Design for escalation, not silent judgment calls. The agent hit a genuine judgment call here: the site said 8PM, the calendar said 9PM, and nobody had told it which one was the mistake. It picked a reasonable default — trust the side that’s already gone out to real people via calendar invites, not the marketing copy — made the fix, and then flagged the call explicitly in a comment instead of just deciding and staying quiet about it. That’s the behavior I designed the prompt around, and it’s the difference between an agent you can leave running and one you have to double-check every morning.


The Part I Didn’t Expect

I went in trying to answer one narrow question — does GitHub access need a separate connector like Linear did. It didn’t; it was already there, auto-attached the same way. That was a fifteen-minute answer.

What I didn’t expect was that the very first unattended-style run would surface a real, live bug that had been sitting on three public pages for weeks. Not a synthetic test case. An actual hour-long discrepancy between what my company promised and what people would actually see when they joined the call.

I don’t think that’s a coincidence, and I don’t think it’s luck. It’s what happens when you build the loop to actually do the work — read the ticket, then go verify the ticket’s premise against the real data — instead of building it to look busy. The verification instinct I applied to the agent’s access is the same instinct I told the agent to apply to the ticket. Trust but verify, at every layer.


What You Can Do Today

  1. Before you schedule any agent loop, trigger it once by hand and read the raw log — not the summary. Fifteen minutes now beats finding out at 5AM that nothing actually happened.
  2. Assume new integrations need explicit authorization until you’ve proven otherwise. Some do, some come for free. Don’t guess either way — check.
  3. Write escalation into the prompt, not just permission. An agent that can push code needs an explicit instruction for what to do when it hits a call only a human should make — and it needs somewhere to put that flag where you’ll actually see it.

The routine runs again tomorrow at 5AM. I already know what its access looks like, because I watched it use every piece of it once, on purpose, before I let it run unattended.


This is Part 1 of a series. Part 2 covers the actual prompt and the guardrails — the cap on open PRs, the never-merge-its-own-PR rule, how it’s told to handle judgment calls — that make it safe to leave running. Part 3 will cover the real results once there are a few mornings of data to report.

Tags: #claude-code#automation#ai-agents#workflow#security-engineer

Comments

Loading comments...