Agentic AI in marketing: why most projects fail, and what a working one looks like

The short answer
Gartner expects more than 40 percent of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
For marketing teams the second one bites hardest, and it has a specific shape. The agent produces a finding, and a person still does all the work.
We have watched this from an unusual angle. Six separate reviews of AI search tooling for this blog reached the same verdict on every product assessed, including ours where it applies. They diagnose. They do not fix. That gap is the business-value problem in concrete form.
Key takeaways
- The failure rate is a design problem rather than a model problem. Gartner names costs, unclear value and weak risk controls. A better model solves none of them.
- “Agent washing” is common. Gartner’s own term for chatbots relabelled as agents. A useful test is whether anything changes in a system you own when the agent finishes.
- Adoption is early. Roughly 17 percent of organisations had deployed AI agents at the time of that research. More than 60 percent expected to within two years.
- Bounded agents beat one broad agent. A narrow brief is testable, correctable and cheap to retire. A general-purpose assistant is none of those.
- The human gate is the point rather than a limitation. Review capacity sets the pace in expert categories. Pretending otherwise is how projects lose internal trust.
Where this comes from. Pepper has run organic growth for more than 250 enterprises across 8 years, and tracks over 10 million prompts on every major engine. We build agents for this work, and we have retired several that did not earn their place. That experience shapes the view here, along with the client reviews we sit in every week and the conversations we have at the events we run.
What is agentic AI in marketing?
Agentic AI in marketing is software that takes an objective, plans the steps, then uses tools and data to carry them out. It produces a result you can act on without reassembling it yourself.
The word does a lot of work, so three distinctions are worth drawing.
- A chatbot answers. An agent acts. If the output is text you then implement yourself, you have an assistant. Gartner calls relabelling one as the other “agent washing”, and it is the most common thing in the category.
- An agent is bounded. A platform is not. Useful agents have a narrow brief, a defined data scope and a testable output. “Marketing AI” is a category, not an agent.
- Autonomy is a dial rather than a switch. Most valuable marketing work sits between fully manual and fully autonomous, with a human approving at one specific point. Brand visibility work is a good example, since the earned half depends on people outside your company entirely.
That middle setting is where nearly all the real value sits today, and it is the least exciting thing to demo. We apply the same test to measurement tools in our review of the tooling market.
See where you show up. Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews. Your team can log in, connect Search Console and GA4, manage the prompt set and run your own agents in the Agent Atlas. A growth team works the same account alongside you. Book a growth audit or see where you show up.
Why most agentic AI projects fail
The numbers are worth stating precisely, along with their date.
Gartner published this prediction on 25 June 2025, attributed to Senior Director Analyst Anushree Verma. The stated basis includes a January 2025 poll of 3,412 webinar attendees on investment posture. We apply a 2026-only rule to research in this field, and this is a deliberate exception. A forward-looking prediction with a 2027 horizon has not expired the way a current-state measurement would. We use it as a statement about direction rather than a current reading.

Three causes are named. In marketing, each has a recognisable form.

- Escalating costs. Agents re-running expensive reasoning over the same unchanged data. The fix is boring. Cache, narrow the scope, and use ordinary code for anything deterministic.
- Unclear business value. The one we see most often. The agent writes a recommendation, a person still does all the work, and nothing downstream changed.
- Inadequate risk controls. No grounding, no human gate, no audit trail. That ends a project outright in regulated categories.
That middle failure is worth dwelling on, because it is measurable from outside. We reviewed the AI search tooling market six times for this blog. Free tools, paid trackers, technical audit tools, full stacks. Every review ended with the same line in the comparison table: diagnoses, never fixes. A dashboard telling you that you are absent from ChatGPT has produced a finding rather than an outcome.
What separates the projects that survive
Nine principles used to sit on this page. In practice they collapse into four that do the load-bearing work.
Each layer depends on the one beneath it. Most failed projects we see have skipped straight to the top.

- Bounded agents, grouped by job. One researches a topic cluster. Another drafts a brief. Another checks claims against sources. Each has a clear remit, so each can be tested and retired on its own. A single broad assistant cannot be evaluated at all, because nothing defines it working.
- Engineered context, not clever prompts. What the agent can see decides output quality more than how the request is phrased. That means real connections to your analytics, your search data and your content system, rather than a text box and hope. The same principle governs written pages, as the citation evidence shows.
- Grounding and a claims check. Every factual claim traced to a source, and an explicit pass that fails the output when it cannot be. This is the control separating useful from dangerous in expert categories, and it maps directly onto how engines judge credibility.
- A human gate at the decision point. Not a rubber stamp at the end. A named person approves the thing carrying risk, with everything before it automated and everything after it executed. In healthcare work that gate is a clinician, and it sets the pace of the whole programme.
The fourth is where most vendors quietly stop, and where the value actually is.
Where Pepper’s approach sits, described accurately
We build in this category, so treat what follows as a worked example rather than a neutral survey.
Our agent layer is the chat interface to the GEO platform. You ask a question in plain English. It reasons across GEO tracking, Search Console, Analytics and your CMS in one pass, shows its working, and gives a straight read rather than a dashboard to interpret.
What it does with the result, stated as the product actually behaves today:
- Output lands in a Sheet. That Sheet is a display method created once a run finishes, rather than a live view of the run. A single run may produce more than one Sheet or document, and typically produces a Sheet once there are five or more outputs.
- Artifacts are agent-generated. They are not user-editable and not interactive, so there is no drilling into a row.
- Documents export as PDF or HTML. Not DOCX.
- CMS access is asymmetric in the current version. Reading works across all three supported systems. Writing back is WordPress-only.
- Moving artifacts between projects is a human job. The agent does not do it.
Why we list the limits. A demo could gloss over every one of those. Gartner’s own framing for this category is “agent washing”, so a specific account of what a product does not do beats another capability list.
Our methodology: how we weight an agentic approach
| Factor | Weight | Why it carries this much |
|---|---|---|
| Does it change something in a system you own | 35% | The difference between an agent and an assistant, and the failure mode Gartner calls unclear business value |
| Bounded, testable agent briefs | 25% | Narrow remits can be evaluated and retired. Broad ones cannot be judged at all |
| Grounding and claim traceability | 25% | The control that keeps expert categories safe, and the one most often missing |
| Human gate at the risk point | 15% | Review capacity sets the real pace in regulated work |
Agentic AI marketing approaches at a glance
| Approach | What it delivers | Typical 2026 cost | Who it suits | Where it falls short |
|---|---|---|---|---|
| A general AI assistant | Drafts and answers, in a chat window | Roughly $20 to $30 per seat a month | Individual productivity | Nothing downstream changes. You implement everything |
| Point AI features in existing tools | Narrow help inside one product | Bundled, or a small add-on | Teams already committed to that tool | Siloed per tool, and no cross-source reasoning |
| A visibility or analytics platform | Findings and dashboards | Roughly $99 to $400 a month | Teams with capacity to act | Diagnoses, never fixes. The work still lands on you |
| Building agents in-house | Exactly your workflow | Engineering time, ongoing | Companies with an AI engineering function | This is the 40 percent Gartner is describing |
| Bounded agents plus a human gate | Findings, plus execution in your systems | Varies with scope | Teams that need the work done, not described | Capped by whoever approves the risky step |
| Platform plus a growth team | Tracking, plus people and agents doing the work | Custom | Teams without internal capacity | Still gated by your reviewers, as it should be |
How to choose an agentic AI approach
One question sorts most of this, and it is not about the model.
The scorecard
Score any agentic proposal out of 100.
| Factor | Weight | How to score it honestly |
|---|---|---|
| Something changes in a system you own | 30% | You can name the system and the change. A document is not a change |
| Each agent has a brief you can read | 25% | Someone can state what one agent does and what would count as it failing |
| Claims are traceable to sources | 20% | You can click from an assertion to where it came from |
| A named human approves the risky step | 15% | The gate is explicit and a person owns it |
| The vendor states what it cannot do | 10% | A specific limits list exists, offered rather than extracted |
The live test, over 90 days
Two halves, both running over 90 days.
First, pick one workflow your team genuinely repeats, with a measurable output. Publishing a brief, refreshing a page, building a report. Run it manually for a month and record two things: hours spent, and how many steps a person had to redo. Then run it with the agent for two more months, recording the same two numbers.
Second, write 30 prompts you would genuinely put to the agent, in the words your team would use. Real ones, such as “which of our pages lost AI Overview visibility last month”, “draft a brief for the query we are losing to a competitor”, “update the pricing table on our comparison page” and “tell me which sources are cited instead of us”.
Run all 30 prompts and score each on one binary: did it change something in a system you own, or did it hand you a document? That ratio is the honest measure of whether you bought an agent or an assistant.
Three readings is the shortest honest window. One is a snapshot, two could be noise, three shows direction. Keep both the workflow and the 30 prompts fixed, because changing either resets the comparison.
Watch the redo count rather than the hours. An agent that halves the time and doubles the rework has not helped, and hours-saved reporting will hide that completely.

Weak approach versus strong approach
The weaker approach: buy a broad assistant, give the team access, then expect productivity to appear. Nothing is bounded, so nothing can be tested. After two quarters nobody can say whether it worked.
The stronger approach: pick one repeated workflow and automate everything up to the decision. Put a named person on that decision, then execute the rest into the system holding the work.
The weak version buys capability. The strong version removes a specific task from a specific person, which is the only one of the two that shows up in a budget.
Red flags
- A demo where the output is a document. If the run ends in text you must implement, that is an assistant.
- No stated limits. Gartner describes this category with the phrase “agent washing”. A vendor naming nothing it cannot do is selling exactly that.
- Accuracy percentages with no method. Nobody in this category publishes an independently checkable one, including us.
- Autonomy as the headline. The gate is the feature in expert categories. Removing it is how projects lose trust.
- One agent that does everything. Untestable by construction, and the first thing to be cancelled.
- Hours-saved as the only metric. It hides rework, which is where agent value usually leaks.
Five questions worth asking
- “What changes in a system we own when this finishes?” The single most useful question here.
- “What can it not do?” A specific answer is a strong signal on its own.
- “Where does a human approve, and who is it?” Vague answers mean no gate.
- “How do we trace a claim back to its source?” If you cannot, it is not usable in expert work.
- “Which agent would you retire first?” Teams that build seriously can answer this immediately.
Reduced to one principle: an agent that ends in a recommendation is an assistant with a longer runtime.
One closing note that costs us something. Most marketing teams do not need an agentic programme. They need one repeated workflow removed, and a good partner will scope that rather than sell a platform, ours included in the general point that platforms are easier to sell than a single fixed workflow.
What does this cost?
| What you are buying | Typical 2026 cost | What it covers |
|---|---|---|
| Baselining one workflow by hand | Free, a month of light tracking | Hours and redo count, the only honest before-picture |
| A general assistant seat | Roughly $20 to $30 per user a month | Drafting and answers, with implementation still on you |
| A visibility or analytics platform | Roughly $99 to $400 a month | Findings across engines, with the work still on you |
| Building in-house | Engineering time, ongoing | Full control, and the failure rate Gartner describes |
| Bounded agents with execution | Varies with scope | Findings, plus changes made in your systems |
| Platform plus a growth team | Custom | Pepper: tracking, plus the people and agents doing the work |
The first row is free, and without it no later row can be judged.
What nobody should promise you
A published accuracy or task-completion rate you can verify. Nobody in this category has one that is independently checkable, and that includes us.
Full autonomy in an expert or regulated category. The human gate is the control that makes the rest usable.
That a better model fixes a failing project. Gartner’s three named causes are cost, value and controls. None is a model problem.
Hours saved as proof of value. It is the metric that most reliably hides rework.
When you do not need agentic AI
The answer that costs us the sale. If you cannot name one workflow your team repeats often enough to measure, you do not need agentic AI yet and we would tell you not to buy it.
Agents pay for themselves on repetition. Without a repeated workflow there is no baseline, no way to separate improvement from noise, and no defensible answer when someone asks what it bought you. Spend a month recording where the hours actually go. That costs nothing, and it is the same first step we would take on your account.
Frequently asked questions
What is agentic AI in marketing?
Software that takes an objective, plans the steps, then uses your tools and data to carry them out. The distinction from a chatbot is that something changes in a system you own, rather than arriving as text for you to implement.
Why do most agentic AI projects fail?
Gartner expects over 40 percent to be cancelled by the end of 2027, naming escalating costs, unclear business value and inadequate risk controls. The middle one dominates in marketing, because the agent produces a finding and a person still does the work.
What is agent washing?
Gartner’s term for relabelling a chatbot or a scripted automation as an agent. The practical test is simple. Ask what changes in a system you own when the run finishes, then listen for whether anything does.
How many companies actually use AI agents?
Roughly 17 percent of organisations had deployed them at the time of Gartner’s research, while more than 60 percent expected to within two years. Adoption is early enough that most published claims describe pilots rather than production.
Should agents run without human review?
Not in expert or regulated categories. The human gate at the decision point keeps the output usable. Removing it is a common route to losing internal trust in the whole programme.
Is one broad AI agent better than several narrow ones?
Narrow is better, mainly because you can judge it. A bounded agent has a brief you test against, and retire if it underperforms. A general-purpose assistant has no definition of working or failing.
How do I measure whether an agent is working?
Baseline one repeated workflow by hand for a month, recording hours and how many steps a person had to redo. Then track the same two numbers with the agent running. Watch the redo count, because hours-saved reporting hides rework.
What should an agentic AI vendor be able to tell me?
What changes in your systems when a run completes, what the product cannot do, where a human approves, and how a claim traces to its source. A vague answer on any of the four is the useful signal.
Where to go next
Ask the one question that sorts this category: when the run finishes, what changed in a system you own?
If the answer is a document, you are looking at an assistant. That may still be worth buying, and it should be priced and judged as one.
Learn AI Search · See the Agent Atlas · Book a growth audit
Sources and further reading
- Gartner. “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027”, press release, 25 June 2025. Attributed to Senior Director Analyst Anushree Verma. Causes named: escalating costs, unclear business value and inadequate risk controls. Stated basis includes a January 2025 poll of 3,412 webinar attendees on investment posture. Also the source of the “agent washing” framing and the adoption figures of roughly 17 percent deployed against more than 60 percent expecting to deploy within two years. Dated 2025 and used here deliberately, as a forward-looking prediction with a 2027 horizon rather than a current-state measurement. Link
- Pepper’s own category reviews for this blog, covering free AI search tools, paid citation trackers, AI visibility trackers, technical audit tooling and full stacks. Every review reached the same conclusion on the products assessed: they diagnose rather than fix. See the AI SEO tool stack, citation tracking tools and the free AI visibility audit.
- Google Search Central. AI search optimisation guidance, 15 May 2026. Warns against providers guaranteeing rankings.
- Deliberately excluded: any agent accuracy, task-completion or hours-saved figure, from us or from any vendor. None is independently checkable.
Latest Blogs
Choosing a GEO agency for mid-market B2B used to be a shortlisting problem. It is now a procurement problem, because the published prices have largely gone. Of four agencies publishing a GEO-specific figure in September 2026, two had withdrawn their pricing pages by early October and one had moved domains. Exactly one still publishes a number a mid-market buyer can act on. So the useful question is no longer which agency is best in the abstract, but how to compare three quotes when only one of them arrived with a method attached.
Third-party sources and AI citations are tightly linked, and the link stops short of where most plans assume. Off-page work decides whether an engine retrieves and cites you. It does not decide whether the answer names you, and 61.7% of brand appearances are citations with no name attached. The two largest 2026 datasets also disagree about how much of the citation surface is third-party at all, one saying 84% and the other putting the brand bucket above half. That disagreement is definitional rather than factual, and it decides where a budget goes.
Why is my competitor showing up in ChatGPT and not me? Before accepting the premise, check it. Across 3,981 brand appearances studied in 2026, 61.7% were citations with no brand name in the answer and only 13.2% produced both. So there are three states that look identical from where you are sitting: genuinely absent, present as an unnamed source, and named less often than a rival. Each has a different cause and a different fix, and working on the wrong one is the most common way this gets expensive.