My AI Agent Opened 9 Pull Requests for One Issue: How I Tamed It and Cut Its Token Bill 140x
· AI & ML · By Zeeshan Ahmad
A few weeks ago I wired an autonomous coding agent (OpenClaw) to one of my Laravel projects. The idea was simple: label a GitHub issue agent:run, and the agent picks it up, writes the code, and opens a pull request for me to review.
Then I opened the repo's pull request tab and found this:
- 9 pull requests for a single issue, each one a fresh attempt at the same feature.
- PRs for issues that had already been implemented and merged weeks earlier.
- Several PRs targeting a branch called
main, in a repo whose real branch ismaster, with diffs of 43,000+ lines because the agent had invented its own base branch. - An issue the agent had closed itself with "Implementation verified, all acceptance criteria met", although nothing had been merged.
And every one of those PRs was opened under my GitHub account.
This post is the debugging story, the fixes, and the pipeline I ended up with, which now does the same job for a fraction of the tokens.
#Part 1: Why Was It Duplicating Everything?
The agent ran on a cron job every 5 minutes with a prompt that boiled down to:
That prompt has no memory. Nothing in it says "I'm already working on this":Find open issues labeled
agent:run. For each one, start a session and complete the implementation.
every 5 min ──► list open agent:run issues ──► issue #6 is still open
│
▼
new branch, new code, new PR (again)
- The issue stays open while a PR is waiting for review, because only a merge closes it. So every poll sees the same issue as new work.
- No check for existing PRs or branches. Each run started from scratch.
- Runs overlapped. A coding session took far longer than the 5-minute interval.
- Weaker models freelance. On a cheaper model the agent decided
mainmust be the base branch, re-did closed issues, and "verified" its own work.
#The fix: put the state on GitHub, not in the prompt
I turned the label itself into a small state machine:
agent:run ──claim──► agent:in-progress ──PR opened──► agent:pr-open
│
└──error──► agent:blocked (with a comment explaining why)
The agent claims an issue by swapping the label before doing any work. The poller only ever looks at agent:run, so a claimed issue becomes invisible to the next run. On top of that:
- One deterministic branch per issue:
agent/issue-<N>. - Before opening a PR, check whether one already exists for that branch. If so, update it instead.
--base master, always. Never createmain.- The agent may never close issues, merge, or approve. A human merges, and the merge closes the issue.
I tested it with a throwaway issue: one claim, one branch, one PR. A second run found nothing to do and created no duplicate.
#Part 2: The Agent Had Far Too Much Access
While debugging I SSH'd into the server and found three things that worried me more than the duplicates.
1. Commits under a mystery name. The "author" on the agent's commits came from the server user's global ~/.gitconfig, set by hand at some point and inherited by every process on the box.
2. A GitHub login with access to everything. The GitHub CLI on the server was logged in as my personal account. The prompt said "only touch this repo", but the token could read and write every repository I own.
3. Its workspace was production. This one made me sit up. The agent's working directory was a symlink to the folder a production Docker container mounts read-write as its live code. Every edit the agent made there was live on that deployment instantly, with no PR and no review.
#Least privilege, applied properly
| Before | After |
|---|---|
Runs as the same user that owns the deploys (passwordless sudo) | Dedicated Linux user ocbot: no sudo, no Docker group, private home |
gh logged in as my account, with access to all repos | Fine-grained token scoped to one repository (contents, issues, PRs only) |
| Git identity set globally for every process | Identity + credential helper set in that one clone's .git/config only |
| Workspace = production code folder | Private workspace; code changes only in a dedicated clone |
Two small tricks made "local to the project" actually enforced rather than hoped for:
# Globally: no identity at all, and refuse to commit without a repo-local one
git config --global user.useConfigOnly true
# gh never stores a login; a wrapper injects the repo-scoped token per call
GH_TOKEN="$(cat ~/.secrets/repo-token)" GH_REPO=owner/repo exec /usr/bin/gh "$@"
Then I verified it the only way that counts, by trying to break it. Pushing to the target repo worked. Reading another private repo returned 404/403. Committing outside the clone failed with "Author identity unknown".
#Part 3: Where Were All the Tokens Going?
I'd seen some eye-watering token counts before, and my first instinct was "the agent is stuffing the whole codebase into the prompt". So I pulled the run history and measured instead of guessing:
| What happened | Tokens |
|---|---|
| A run where there was nothing to do (average of 12) | ~171,000 |
| Creating a 3-line Markdown file from a test issue | 709,263 |
The 3-line file is the interesting one. Reading the run's transcript, the cost came from the shape of the loop, not from reading code:
- Every LLM call re-sends the whole conversation, which starts with a ~32k-token preamble: the agent's personality files, tool definitions and skills list.
- The model ran one shell command per call:
git fetch,git switch,git commit,gh pr createand so on. That's 23 calls for a trivial task. - 23 × ~30k ≈ the 700k I was billed.
The root causes, in order of cost:
1. The LLM was woken up even when there was no work. Every poll cost ~171k tokens to say "nothing to do". 2. A heavy fixed preamble, paid again on every turn. 3. Using a language model for mechanical steps that a shell script does deterministically and for free. 4. Exploring the codebase to find the right files. This is the one everyone talks about, and it was actually last on the list.
#Part 4: Let a Script Do the Plumbing, Let the LLM Write Code
I rebuilt the pipeline so the language model does exactly one job: edit code. Everything else is plain, testable shell.
cron (every 30 min) ─► shell worker (no LLM)
1. any agent:run issue? no → exit 0 tokens
2. claim label, skip if PR exists, branch from master 0 tokens
3. refresh code graph (AST only) 0 tokens
4. pick the ~10 files the issue probably needs 0 tokens
5. openclaw agent --agent coder ◄── the ONLY LLM step
6. guardrails → commit → push → PR → label 0 tokens
#Picking the right files without an LLM
For step 4 I used Graphify, which parses the codebase with tree-sitter into a queryable knowledge graph. For code it is fully local and deterministic, with no LLM calls. On this repo it built a graph of ~8,000 symbols and ~13,000 links in under a minute.
Graphify alone wasn't enough for Laravel, though. A lot of Laravel's wiring is strings, not code references:
Route::group(['prefix' => 'vendor', 'namespace' => 'Nero\Ecomm\...\Admin'], function () {
Route::resource('products', 'Products\ProductController');
});
// and inside the controller:
$this->viewPath = getEcommViewTagHelper() . 'admin.ecomm.product.';
return view($this->viewPath . 'index');
An AST graph can't see that /vendor/products leads to that controller, or that 'admin.ecomm.product.' is a folder of Blade templates. So the file picker combines three signals:
| Signal | What it catches |
|---|---|
| Laravel route chain: URL in the issue → route → controller → views | Pages and endpoints, including Route::resource, group namespaces and runtime view paths |
| Graphify query over the code graph | Related models, services, base classes |
| Keyword search (ripgrep), only when the match is strong | Exact identifiers mentioned in the issue |
Files that only match weakly are dropped, so when an issue doesn't map to existing code the agent gets an empty list instead of noise. On a real feature issue (a vendor product-management redesign), the picker found the right route, controller, layout, views and model in about 5 seconds.
#A lean "coder" agent
The editing agent is stripped to the bone:
- a minimal instruction file: what to change, what never to touch, don't run git;
- no skills list, coding tools only (no browser, cron or messaging), no hidden "thinking";
- any single tool result capped at 12k characters, so one giant file can't flood the context.
Its fixed preamble went from ~32,000 tokens to ~2,800.
#Guardrails the model can't talk its way past
After the agent finishes, the script checks the diff before anything is committed:
- changes to
vendor/,public/,storage/,.envor SQL dumps are reverted; - more than 25 files or 1,500 changed lines → no PR, the issue is labeled
agent:blocked; - the PR body includes the agent's own summary, the diff stats, the files it was given, and the token count, so cost is visible on every review.
#The Results
Same task, same model, before and after:
| Before | After | |
|---|---|---|
| Run with nothing to do | ~171,000 tokens | 0 (1.2 s, no model call) |
| Test issue → pull request | 709,263 tokens, 23 LLM calls | 4,974 tokens |
| Fixed context per LLM call | ~32,000 tokens | ~2,800 tokens |
| Duplicate PRs per issue | up to 9 | 1 (verified on a second run) |
| What the agent can reach | every repo I own, plus production | one repo, one clone |
That's roughly 140x cheaper on the test task, and idle polling went from burning millions of tokens a day to free.
#Lessons Learned
#1. Measure before you optimize
"The agent is sending the whole repo to the LLM" felt obviously true. The run history said otherwise: most of the waste was idle polling and a chatty loop. Graphify still earned its place, just not as the first fix.#2. Keep state outside the model
Labels, branch names and PR checks live on GitHub, where every run can see them. A prompt can't forget something it never had to remember.#3. Give the LLM the smallest job possible
Deterministic steps belong in code: picking, claiming, branching, committing, opening PRs. The model is expensive, slow and creative. Spend it only where creativity is needed.#4. Least privilege isn't optional for agents
A dedicated user, a token scoped to one repo, git config local to one clone, and a workspace nowhere near production. Then try to break it, and confirm that it fails the way you expect.The agent now runs every 30 minutes. Most of the time it costs nothing. When I label an issue, I get one clean pull request, with its token bill in the description.