My AI Agent Opened 9 Pull Requests for One Issue: How I Tamed It and Cut Its Token Bill 140x

· AI & ML · By Zeeshan Ahmad

A few weeks ago I wired an autonomous coding agent (OpenClaw) to one of my Laravel projects. The idea was simple: label a GitHub issue agent:run, and the agent picks it up, writes the code, and opens a pull request for me to review.

Then I opened the repo's pull request tab and found this:

  • 9 pull requests for a single issue, each one a fresh attempt at the same feature.
  • PRs for issues that had already been implemented and merged weeks earlier.
  • Several PRs targeting a branch called main, in a repo whose real branch is master, with diffs of 43,000+ lines because the agent had invented its own base branch.
  • An issue the agent had closed itself with "Implementation verified, all acceptance criteria met", although nothing had been merged.

And every one of those PRs was opened under my GitHub account.

This post is the debugging story, the fixes, and the pipeline I ended up with, which now does the same job for a fraction of the tokens.


#Part 1: Why Was It Duplicating Everything?

The agent ran on a cron job every 5 minutes with a prompt that boiled down to:

Find open issues labeled agent:run. For each one, start a session and complete the implementation.

That prompt has no memory. Nothing in it says "I'm already working on this":
text
 every 5 min ──► list open agent:run issues ──► issue #6 is still open
                                                  │
                                                  ▼
                                   new branch, new code, new PR  (again)
  • The issue stays open while a PR is waiting for review, because only a merge closes it. So every poll sees the same issue as new work.
  • No check for existing PRs or branches. Each run started from scratch.
  • Runs overlapped. A coding session took far longer than the 5-minute interval.
  • Weaker models freelance. On a cheaper model the agent decided main must be the base branch, re-did closed issues, and "verified" its own work.
IMPORTANT
An autonomous agent's prompt is not a guardrail. If the safety of your workflow depends on the model remembering or choosing to behave, it will eventually not behave. State has to live somewhere the model can't forget it.

#The fix: put the state on GitHub, not in the prompt

I turned the label itself into a small state machine:

text
agent:run ──claim──► agent:in-progress ──PR opened──► agent:pr-open
                            │
                            └──error──► agent:blocked  (with a comment explaining why)

The agent claims an issue by swapping the label before doing any work. The poller only ever looks at agent:run, so a claimed issue becomes invisible to the next run. On top of that:

  • One deterministic branch per issue: agent/issue-<N>.
  • Before opening a PR, check whether one already exists for that branch. If so, update it instead.
  • --base master, always. Never create main.
  • The agent may never close issues, merge, or approve. A human merges, and the merge closes the issue.

I tested it with a throwaway issue: one claim, one branch, one PR. A second run found nothing to do and created no duplicate.


#Part 2: The Agent Had Far Too Much Access

While debugging I SSH'd into the server and found three things that worried me more than the duplicates.

1. Commits under a mystery name. The "author" on the agent's commits came from the server user's global ~/.gitconfig, set by hand at some point and inherited by every process on the box.

2. A GitHub login with access to everything. The GitHub CLI on the server was logged in as my personal account. The prompt said "only touch this repo", but the token could read and write every repository I own.

3. Its workspace was production. This one made me sit up. The agent's working directory was a symlink to the folder a production Docker container mounts read-write as its live code. Every edit the agent made there was live on that deployment instantly, with no PR and no review.

#Least privilege, applied properly

BeforeAfter
Runs as the same user that owns the deploys (passwordless sudo)Dedicated Linux user ocbot: no sudo, no Docker group, private home
gh logged in as my account, with access to all reposFine-grained token scoped to one repository (contents, issues, PRs only)
Git identity set globally for every processIdentity + credential helper set in that one clone's .git/config only
Workspace = production code folderPrivate workspace; code changes only in a dedicated clone

Two small tricks made "local to the project" actually enforced rather than hoped for:

bash
# Globally: no identity at all, and refuse to commit without a repo-local one
git config --global user.useConfigOnly true

# gh never stores a login; a wrapper injects the repo-scoped token per call
GH_TOKEN="$(cat ~/.secrets/repo-token)" GH_REPO=owner/repo exec /usr/bin/gh "$@"

Then I verified it the only way that counts, by trying to break it. Pushing to the target repo worked. Reading another private repo returned 404/403. Committing outside the clone failed with "Author identity unknown".

TIP
If your agent pushes as you, give it a token that can only touch the one repo it works on. Everything it does still shows up as you on GitHub, but the damage it can do is limited to that repo.

#Part 3: Where Were All the Tokens Going?

I'd seen some eye-watering token counts before, and my first instinct was "the agent is stuffing the whole codebase into the prompt". So I pulled the run history and measured instead of guessing:

What happenedTokens
A run where there was nothing to do (average of 12)~171,000
Creating a 3-line Markdown file from a test issue709,263

The 3-line file is the interesting one. Reading the run's transcript, the cost came from the shape of the loop, not from reading code:

  • Every LLM call re-sends the whole conversation, which starts with a ~32k-token preamble: the agent's personality files, tool definitions and skills list.
  • The model ran one shell command per call: git fetch, git switch, git commit, gh pr create and so on. That's 23 calls for a trivial task.
  • 23 × ~30k ≈ the 700k I was billed.

The root causes, in order of cost:

1. The LLM was woken up even when there was no work. Every poll cost ~171k tokens to say "nothing to do". 2. A heavy fixed preamble, paid again on every turn. 3. Using a language model for mechanical steps that a shell script does deterministically and for free. 4. Exploring the codebase to find the right files. This is the one everyone talks about, and it was actually last on the list.


#Part 4: Let a Script Do the Plumbing, Let the LLM Write Code

I rebuilt the pipeline so the language model does exactly one job: edit code. Everything else is plain, testable shell.

text
cron (every 30 min) ─► shell worker  (no LLM)
  1. any agent:run issue?            no → exit           0 tokens
  2. claim label, skip if PR exists, branch from master  0 tokens
  3. refresh code graph (AST only)                       0 tokens
  4. pick the ~10 files the issue probably needs         0 tokens
  5. openclaw agent --agent coder  ◄── the ONLY LLM step
  6. guardrails → commit → push → PR → label             0 tokens

#Picking the right files without an LLM

For step 4 I used Graphify, which parses the codebase with tree-sitter into a queryable knowledge graph. For code it is fully local and deterministic, with no LLM calls. On this repo it built a graph of ~8,000 symbols and ~13,000 links in under a minute.

Graphify alone wasn't enough for Laravel, though. A lot of Laravel's wiring is strings, not code references:

php
Route::group(['prefix' => 'vendor', 'namespace' => 'Nero\Ecomm\...\Admin'], function () {
    Route::resource('products', 'Products\ProductController');
});
// and inside the controller:
$this->viewPath = getEcommViewTagHelper() . 'admin.ecomm.product.';
return view($this->viewPath . 'index');

An AST graph can't see that /vendor/products leads to that controller, or that 'admin.ecomm.product.' is a folder of Blade templates. So the file picker combines three signals:

SignalWhat it catches
Laravel route chain: URL in the issue → route → controller → viewsPages and endpoints, including Route::resource, group namespaces and runtime view paths
Graphify query over the code graphRelated models, services, base classes
Keyword search (ripgrep), only when the match is strongExact identifiers mentioned in the issue

Files that only match weakly are dropped, so when an issue doesn't map to existing code the agent gets an empty list instead of noise. On a real feature issue (a vendor product-management redesign), the picker found the right route, controller, layout, views and model in about 5 seconds.

#A lean "coder" agent

The editing agent is stripped to the bone:

  • a minimal instruction file: what to change, what never to touch, don't run git;
  • no skills list, coding tools only (no browser, cron or messaging), no hidden "thinking";
  • any single tool result capped at 12k characters, so one giant file can't flood the context.

Its fixed preamble went from ~32,000 tokens to ~2,800.

#Guardrails the model can't talk its way past

After the agent finishes, the script checks the diff before anything is committed:

  • changes to vendor/, public/, storage/, .env or SQL dumps are reverted;
  • more than 25 files or 1,500 changed lines → no PR, the issue is labeled agent:blocked;
  • the PR body includes the agent's own summary, the diff stats, the files it was given, and the token count, so cost is visible on every review.

#The Results

Same task, same model, before and after:

BeforeAfter
Run with nothing to do~171,000 tokens0 (1.2 s, no model call)
Test issue → pull request709,263 tokens, 23 LLM calls4,974 tokens
Fixed context per LLM call~32,000 tokens~2,800 tokens
Duplicate PRs per issueup to 91 (verified on a second run)
What the agent can reachevery repo I own, plus productionone repo, one clone

That's roughly 140x cheaper on the test task, and idle polling went from burning millions of tokens a day to free.

NOTE
The test issue was deliberately trivial. Real features will cost more because the agent has to read and change real code. But the overhead that used to dominate (idle polling, the heavy preamble, LLM-driven git plumbing) is gone, and every PR now reports its own cost.

#Lessons Learned

#1. Measure before you optimize

"The agent is sending the whole repo to the LLM" felt obviously true. The run history said otherwise: most of the waste was idle polling and a chatty loop. Graphify still earned its place, just not as the first fix.

#2. Keep state outside the model

Labels, branch names and PR checks live on GitHub, where every run can see them. A prompt can't forget something it never had to remember.

#3. Give the LLM the smallest job possible

Deterministic steps belong in code: picking, claiming, branching, committing, opening PRs. The model is expensive, slow and creative. Spend it only where creativity is needed.

#4. Least privilege isn't optional for agents

A dedicated user, a token scoped to one repo, git config local to one clone, and a workspace nowhere near production. Then try to break it, and confirm that it fails the way you expect.

The agent now runs every 30 minutes. Most of the time it costs nothing. When I label an issue, I get one clean pull request, with its token bill in the description.

Related articles

Home · All Tools · Blog