Translated from the Japanese original on mureo.jp.
When the subject of handing ad ops to Claude Code comes up, the first question is how far it can go. Read what people actually using it have published, and the delegated range stops in almost the same place. An operator running ¥50 million a month single-handedly at a brand, someone two years into outsourced ad management, an office worker’s personal build, a developer who has kept several tools wired together running for a fortnight. Different industries, different scales, no sign of having compared notes, and yet the line sits in the same spot.
This article sets out where that line is, why it settles there, and what to judge by when you draw your own. The risks of delegating at all and the safety-side design are in Is it safe to hand ad ops to Claude Code?, and what can be automated in the first place is in Can you automate ad ops with Claude Code?.
One caveat: the accounts below are public, but we have not verified the authors’ businesses or their results. No names here, only the shape that is common to all of them.
What gets delegated stops at gathering the numbers
The work actually handed over falls into five kinds.
Scheduled retrieval and notification. Pulling spend and cost per conversion hourly into a chat channel, delivering yesterday’s figures as a morning digest. The point is to open the dashboard less often; one account simply says they stopped opening it.
Transcription across platforms. Collecting several ad platforms plus cart and order data into one sheet every morning. Done by hand it takes thirty minutes to an hour, and it recurs daily. Nothing is being judged, only copied, and a mistake is visible on inspection. As candidates for delegation go, this is the straightforward kind.
Collecting change history. Pulling when budgets moved and which ads were stopped out of the platform’s history page each day, into a record of your own. As below, this is a person filling in for something the platform does not provide.
Producing candidate lists. Aggregating search terms over seven days and keywords over thirty, then producing candidates for negatives and for pausing, with the reasoning attached. What comes out is a list; the decision to apply it stays with a person.
Drafting and checking text. Producing ad copy options, transcribing competitors’ video ads, checking cosmetics and supplement scripts against regulated wording. That last one is the kind of task where a human misses more than they catch.
Line the five up and the shared property is visible. Each leaves its result as text or numbers, so a mistake is apparent if you read it. And none of them, on execution, moves money out of the account.
What stays behind touches money and stopping
Three things are kept: the bidding decision, budget allocation, and whether to stop delivery.
The account from someone two years into outsourced ad management is the most concrete. They ran an experiment handing over the whole set at once, daily reports, figure checks, creative options, upload, even support for budget decisions. Ten hours of building later it did not behave as expected and they went back to doing it by hand. What they delegate now is three things, reports, figure checks and creative options; bidding, budget allocation and delivery decisions they do themselves.
The operator running ¥50 million a month alone hands over the five tasks above and still keeps the bidding decision. In the setup where several tools produce a daily brief, the human review has shrunk to about five minutes, but writes are limited to what has been approved.
So the line does not fall between reading and writing, but between work that moves money and work that does not. Drafting ad copy is writing, yet nothing happens as a result. The character of the task changes the moment it is uploaded and delivery starts.
Why the line settles in the same place
Three reasons, none of them about how mature the technology is.
The first is that the ways back differ. A wrong number in a report is fixed by reading it again. A mistaken bid keeps spending until someone notices. With automated bidding it is worse: restoring the setting restarts the learning phase, so the figures are unsettled for days. Whether the cost of a mistake is recoverable decides whether the task can be handed over.
The second is that correctness is hard to guarantee. The failures listed in the experiment that was rolled back were overwriting unrelated cells, producing wrong figures, and returning a table read out of an image in a form that could not be edited. None of these stops the work. They are failures of the kind where it runs, but the result cannot be trusted. In a report a person notices. Underneath a bidding decision, there is no way to notice.
The third is that there is someone to explain it to. An agency answers to its client, an in-house team to the business. Hand over the decision and the explanation becomes “the agent decided that.” This is a question of where responsibility sits, so improving accuracy does not dissolve it.
Drawing the line costs more than the work
What gets overlooked is the cost of deciding how far to delegate. The record of that rolled-back experiment says, in effect, that drawing the boundary was harder than the work itself.
That rings true. Deciding the range means working through, task by task, what happens when each one fails. Daily work has a settled procedure and runs without thought; designing the boundary takes a judgment every time. The ten hours of building that produced nothing is better read not as a tooling problem but as most of those hours going into trial and error over the range.
Which suggests that handing over a wide range at the start is the slower route. Delegate one of the five kinds, the morning transcription say, and run it for a fortnight. Look at the shape of the failures before adding the next. Designing takes less time once you have material for the decision in hand.
The sticking point is where the platform keeps no record
That “collecting change history” appears among the delegated tasks is worth pausing on, because it is something the dashboard could reasonably provide itself.
Indeed, what the operator running ¥50 million a month named as their sticking point was the absence of any single view of budget-operation history. Hence the job of scraping the history page daily into their own store. Google Ads has a change-history feature, but granularity and retention differ by platform and nothing lines several platforms up side by side.
This shape of task is a high-value candidate for delegation. No judgment enters, it recurs daily, and people do not keep it up by hand. The accumulated history is also the material for working out later why a number moved. Look at outcomes alone, with no record of what changed and when, and the causes cannot be separated. Key-person dependency and the design around it are covered in Why does ad ops end up depending on one person?.
The order to draw your own line in
Apply four questions to each task in turn.
- On failure, does money go out? If it does, do not delegate. If it does not, move on. Retrieval, aggregation, transcription and drafting pass here. Upload passes in part if you separate producing the draft from pushing it live.
- On failure, is it apparent on reading? If it is, delegation is fine. What cannot be noticed is dangerous even when no money moves. Aggregation can be brought inside this condition by building the reconciliation against source values into the same procedure.
- Is the procedure the same every time? Work whose criteria shift cannot be handed over in prose. Where the conditions can be fixed, the criteria can be written out in plain language and passed across. There are people building this way without writing code, explaining the criteria in prose and correcting what comes back.
- Does a record remain? Where what was done cannot be read back afterwards, do not widen the range.
Delegate only what passes all four; leave the rest at proposing candidates. For negative keywords, hand over producing the list and the reasoning, and let a person apply it. That satisfies 2 and 4 at once.
Building the approval step
Keeping decisions in hand is no good if checking them in the dashboard returns you to the original workload. The setups that have lasted narrow where and when the checking happens.
The arrangements that work look alike. Figures and candidates are gathered automatically into one place. The person looks only there. Only what a person has approved gets executed. In one account that check fits into five minutes a day.
A caution: interposing approval does not by itself make things safe. If the reasoning behind a proposal is not legible, people cannot examine it and the approval becomes a formality. Require the reasoning and the underlying numbers alongside every proposal. This is covered in more depth in Is it safe to hand ad ops to Claude Code?.
Frequently asked questions
Is nobody delegating bidding and budgets?
Excluding promotional posts and people selling information products, we found almost no continuing cases. We cannot assert there are none, but within what we searched, no one writing publicly says they have taken the bidding decision away from a human and kept it that way. Not because it is technically impossible, we read, but because the three reasons above are doing the work.
Can you build this without writing code?
There are examples. The operator running ¥50 million a month states plainly that they cannot write code, and describes explaining the criteria for a decision in plain language, having it built, running it, and having it corrected. That said, handing over a wide range at the start also multiplies what needs correcting. Start with one task and add the next once it runs.
Where is the best place to start?
With work that recurs daily, involves no judgment, and shows its mistakes on inspection. Collecting numbers from several platforms into one place meets those conditions best. The effect is legible too, and the gap against doing it by hand widens the longer it runs. Per-platform connection steps are in Google Ads and Meta Ads.
Should you connect over MCP or the API?
It depends on the use. MCP is quicker to connect and sufficient for calling settled operations. Once you need fine-grained conditions or bulk retrieval, what MCP exposes runs short, and there is an account of switching to calling the API directly. Starting on MCP and reconsidering when it runs short is fine.
When should the range be revisited?
When something fails, and when the platform’s behaviour changes. There is no need to review on a schedule, but if people have stopped reading the output of a delegated task, that is a signal to inspect it. Output nobody reads is also output whose correctness nobody is checking.
Summary
Line up the accounts of people running ad ops through Claude Code and the delegated range stops in the same place. Handed over: scheduled retrieval and notification, transcription across platforms, collecting change history, producing candidates for negatives and pausing, drafting and checking text. Kept: bidding, budget allocation, and stopping delivery. The line falls not between reading and writing but between work that moves money and work that does not, for three reasons, how recoverable the failure is, whether correctness can be guaranteed, and whether there is someone to explain it to.
Drawing your own: ask of each task whether money goes out, whether a failure is apparent on reading, whether the procedure can be fixed, and whether a record remains. Delegate what passes all four and leave the rest at proposing candidates. Because deciding the range itself takes time, getting one thing running and then adding is faster than handing over a wide range at the start.