AI is genuinely good at reading an Amazon ad account and should not hold unsupervised authority over it. The dangerous failure isn't a bad decision — it's a system that quietly loses the ability to tell you whether it's making bad ones. Why our AI advises and our rules execute.
We automate the work of Amazon advertising aggressively. What we will not automate is unbounded model judgment. Those are different things, the category sells them as one, and the difference is where your ad spend goes.
What AI is genuinely good at in Amazon advertising
Start with the concession, because nothing after it counts otherwise. AI belongs in Amazon PPC. It is the authority we are arguing about, not the presence.
A model reading ninety days of performance across ten thousand keywords is doing something a person genuinely cannot. It does not get bored at keyword four thousand. It does not have a favourite campaign, or an ad group it stopped opening in March because the numbers there make it feel bad. It will find the converting search term buried three levels down, and it will find it again tomorrow night, without ever deciding it has already looked at that one.
So take the whole premise as read: AI is good at this. Experienced operators are good at this. Amazon advertising is genuinely hard, which is why both of those sentences need to be true at once. Nothing below is an argument that the model is stupid. It is an argument about who holds the authority — and about who gets to decide that, which is not us.
The dangerous failure is not a bad decision
The most dangerous failure is not a system making bad decisions. It is a system that has quietly lost the ability to tell you whether it is making bad decisions.
A bad decision is survivable. It shows up in the numbers, somebody argues about it, it gets reversed. A system that can no longer report on itself is a different animal, because every reading it gives you is reassuring — including the readings it would give you if it were completely broken.
This is also why "explainable AI" is not the answer, and why we are not going to claim it as one. Explanations are cheap and getting cheaper. Any competitor can bolt a why-this-bid panel, some feature attributions and a decision log onto a fully autonomous system, and then say — truthfully — that they are not a black box any more. None of that touches the harder question underneath it: who is checking that the checking works?
Count the separate pieces of software involved. The model is one. The pipeline feeding it is another. The deterministic engine that executes is a third, the record of what it did is a fourth, and the counter reporting on all of it is a fifth — written by the same people, shipped on the same day, carrying the same bugs. Four of those five are not AI at all, which is exactly why calling their failures "AI failures" lets everybody off the hook.
And that layer does not fail loudly. It fails in the shape of a zero, and a zero is the most reassuring number on the page.
The failure mode is not error. It is confidence.
People picture AI risk as a calculator that is occasionally off — a number that looks strange, an answer you would catch. That is not how these systems fail.
A model produces the most plausible output available to it. When the inputs are good, plausible and correct are the same thing. When the inputs are incomplete, stale, or quietly mislabelled somewhere upstream, plausible and correct come apart — and the output is still plausible. It still reads as careful reasoning. It still cites specific numbers. It is, in tone and structure and confidence, indistinguishable from every answer that was right.
A wrong AI recommendation does not look like a wrong answer. It looks like a right answer.
We have watched our own engine produce a bid recommendation that was internally coherent, well-argued, and computed against the wrong keyword's performance. Nothing in the output betrayed it. The reasoning was sound. The arithmetic was correct. The inputs were wrong, and the model had no way to say so, because saying so was never its job.
Better models reduce some of this. A stronger model catches more contradictions, flags more anomalies, and hedges more often where the data is thin — that is a genuine improvement and we buy it every time it ships. But better models reduce some errors; they do not make bad inputs trustworthy. The model is usually not the broken thing. The pipeline feeding it is. A model cannot recover provenance the pipeline never preserved.
Three AI PPC failures where the model was never the problem
Three failures out of our own account. Not one of them was the model being stupid, and not one would have shown up on a dashboard. All three have the same shape: the software carried on reporting normally after the thing it was reporting on had stopped being true.
A counter that reads zero looks exactly like nothing happening. When your monitoring reports zero changes held today, there are two possible worlds: nothing was held, or the counter is broken. From the dashboard they are the same picture. We have lived in the second world — our own instrumentation reported zero on every run for two days while the system was holding changes constantly. Everything was green. Nothing was failing. Nothing was being measured, either.
Every AI tool in this category has a dashboard saying everything is fine. Ask what that dashboard would say if it were broken.
Silence gets read as an answer. Amazon's API does not always answer. Sometimes it returns nothing — not an error you could alert on, just an empty space where a number should be. A system that is not careful treats that emptiness as information. "No suggested bid returned" quietly becomes "Amazon disagrees with your bid", and a decision is being made on a fact nobody ever stated. We shipped exactly this bug. We caught it because we count how often each decision path fires, and the counts moved.
A success response is not proof anything happened. You switch a rule off. The API returns success. Except the spend keeps arriving, because "off" landed on one entity and not the adjacent one, or on a level of the hierarchy that was not the level doing the spending. The confirmation was genuine. It just confirmed something other than what you asked for. An AI operating your account reads that same success response and moves on, permanently satisfied.
Accuracy is the wrong question. Exposure is the right one.
Here is the sentence that ends most sales calls in this category: our AI is right about ninety-five percent of the time. It sounds like an answer. It is not even the right question.
Even a very accurate model becomes dangerous when its mistakes are multiplied across thousands of opportunities to act. A model that is wrong once in twenty is completely fine when it makes twenty decisions and a person reads all of them. The same model, at the same accuracy, is a liability when it gets ten thousand chances to act every night and nobody is required to look. The accuracy did not change. The exposure did.
The relevant question is not the model's average accuracy. It is how many wrong actions can reach your account before a human notices.
That reframing turns a modelling question into an engineering one, and engineering questions have answers you can demand. How many changes go out unreviewed in a single run? How large is any one of them allowed to be? What is the total movement permitted in a day before something stops on its own? How long can a wrong decision sit in the account before it surfaces somewhere a person actually looks? Every one of those has a number. A vendor either knows it or has never thought about it.
An AI explanation is not an audit record
Amazon advertising is not a modelling problem. It is an accountability problem wearing a modelling problem's clothes. When your ACoS goes sideways across a quarter, somebody has to answer four questions: what changed, when, who changed it, and on what evidence.
And a model will happily take part in that conversation. Ask it why, and it will tell you — fluently, plausibly, at whatever length you asked for. Every vendor in this category will demo exactly that, and it is the most persuasive ninety seconds in the pitch. It is also the trap.
A model can explain a decision. That does not mean the explanation is a reproducible record of why the decision occurred.
The explanation is generated after the fact, by the same machinery that produced the decision, and it is optimised to read well rather than to be complete. Ask twice and you can get two defensible reasons for one action. Neither is a lie. Neither is a record. Compare it with a row that says: rule "High ACoS 30d" fired; it saw 47.2% ACoS, 31 clicks, $84.10 spend, 2 orders; bid moved $1.42 to $1.21, a 15% cut inside a $0.35 floor and a $2.00 ceiling; auto-applied; reversible.
Nothing in that row was written to persuade anybody. It is not an account of the decision — it is the decision, with the evidence it stood on. An explanation you cannot reproduce is a story. A row you can reproduce is a record. Only one of them is any use when the question is where the money went.
Yes, our top tier runs unattended. That is the point, not the contradiction.
Anyone still reading has spotted it: we sell a plan that changes your account without asking first. Good — this is the most important part. Distrust is not the opposite of automation. Distrust is the thing that makes automation safe. We spent the engineering on three properties before letting anything run unsupervised, and none of them are glamorous.
Bounded. Every bid a rule moves has a limit you set — a floor under anything that cuts, a ceiling over anything that raises. Not a suggestion handed to the model: an arithmetic limit applied after it, which it does not get to argue with. A recommendation landing outside your bounds is clamped or refused, and we tell you which.
Reversible. Bids and on/off states go back with a button — a recorded prior state, not "you can set it back by hand". Undo restores the setting. It does not restore the spend, and anyone who tells you otherwise is selling. What it buys back is time: bounded limits how far a wrong change can move, reversible limits how long you live with it. Neither is worth much alone.
Inspectable. Every automated action writes a row: what changed, from what to what, which rule fired, and the numbers it fired on. You can read it, disagree with it, and find the specific one that broke things — the only kind of finding that helps at 9pm on a Thursday.
Where that stops, since we are about to tell you to make vendors say it. Automated bid and budget changes retain their prior state and can be reversed. Negatives are the exception, and it is worth being blunt rather than clever: a negative keyword can be deleted, but there is no prior state to restore and the impressions it blocked while it was live are gone. Auto-apply on negatives exists, and it is off until you switch it on. That distinction is the whole product. Being able to do a thing and being pushed into it are different, and our job is to tell you plainly what a switch does and what it will not give back — then get out of the way.
Now look at what those properties have in common. None of them are AI. They are arithmetic, a log table, and an undo stack. They are profoundly boring, and they are the entire reason autonomy is a responsible thing to sell.
The autonomy is not the product. The guardrails are the product. The autonomy is what they buy you.
And the thing actually running unattended is not model judgement. It is rules. A rule is a sentence you wrote: if ACoS is above 40% over 30 days with at least 15 clicks, cut the bid 15%, never below $0.35. You can read that. You can predict it. You can work out by hand what it would have done to last month, and get the same answer every time you check. Try reading a model.
AI proposes. Rules execute. Guardrails enforce. Everything records.
A percentage of spend is a bid-raising incentive bolted to a system nobody has to review
Flat pricing on its own is not much of an argument and we are not going to oversell it. But check before you compare anyone on features: this is not only an agency habit. Software in this category takes a percentage of your ad spend too, and the tool you are weighing us against may well be one of them.
The argument is about a compound. It needs all three parts at once: the vendor earns more when your ad spend rises, the vendor's software can raise your ad spend, and nobody is required to look before it does.
Any one of these on its own is ordinary. Agencies have been paid on a percentage of spend for twenty years, and most of them do not ring you before every bid change either — what they have is a named human on the account and a review you can put in the diary. Somebody answers for the quarter out loud. Autonomous software on a flat fee is fine too: nothing about a fixed monthly price gets more attractive when your bids go up. So the missing ingredient is not approval per action. It is that somewhere in the arrangement there is a person whose job is to look, on a cadence you can name.
Assemble all three and you have something that raises its own revenue, unsupervised, and reports the result to the person paying for it. The intent is almost beside the point. The incentive exists whether anyone designed for it or not, and it does not need a villain to operate.
Six questions to ask any AI Amazon PPC vendor
These are the ones we would find uncomfortable, which is exactly why they are the right ones to ask.
- Show me the last change your system made to my account, and the specific numbers it made that change on.
- What happens when Amazon's API does not answer? Does your system know the difference between "no" and "no reply"?
- Can I set a hard ceiling your AI cannot exceed — and is it applied before the model or after it?
- How do I undo last night?
- Have you evaluated your AI's accuracy against ground truth? What was the number, and when did you last re-run it?
- If your own monitoring broke, what would my dashboard look like?
If the answer to the last one is "that would not happen", you have learned everything you need. It happens. It has happened to us more than once. The only variable that matters is whether anybody is looking.
We are not going to decide for you whether the AI is any good
Everything above is an argument about where a model belongs. Here is the thing we refuse to do with that argument: settle it on your behalf.
Some sellers will read all of this, run the AI recommendations, approve them over coffee every morning and never write a rule. Some will never switch the AI on at all and build twenty rules instead, because they already know their account and they want arithmetic, not opinions. Some will run unattended and read the log on a Friday. Those are all people using this correctly.
You can work out whether the AI suits your account better than we can. That is not humility. It is just true.
Which is why the pitch is not "sit back, our AI handles it". That pitch asks you to take the most consequential thing in your business — what you pay per click, across every product you sell — and hand it to a stranger's software on trust, before you have any evidence. We would not accept that deal. We are not going to offer it to you and call it a feature.
What we sell instead is the equipment and the visibility to make that call yourself. One flat price, $129 a month, whether you spend five hundred a month on ads or fifty thousand: Sponsored Products, Sponsored Brands and Sponsored Display, review request automation, the rules engine, the AI recommendations, per-product profit, the audit log, and undo wherever undo is technically possible. No percentage of your spend. No call to find out the price. No tier where trusting us more is how you unlock the useful parts.
Use the AI. Ignore the AI. Turn on unattended and turn it off again in a fortnight because you did not like what Thursday looked like. All fine. You are running a business, not applying for permission, and the software should be arranged accordingly.
Where we land
We are an AI company. Our recommendation engine is a model. Our assistant is a model. We spend real money on inference every night and it earns its keep by a comfortable margin.
We simply do not confuse "this is extremely good at reading" with "this should have the keys". One of those is a tool. The other is a story the category tells because "we built an undo button and a great deal of careful arithmetic" does not fit on a landing page.
The AI advises. You decide. And when you would rather not decide every single time, then rules you wrote decide, inside limits you set, leaving a record you can read and a button that puts it back where a button can.
The autonomy was never the hard part. The guardrails were. That is not caution about AI — that is seriousness about your money.
Want to see what your account looks like before you decide any of this? The free 30-second PPC audit grades it in a minute, and the pricing page has the number on it — no call required.