Free store scan · Founding Brand Programme: 14 days free Scan your store →

liftable

Guide · AI shoppers

What autonomous agents actually do for a store

“Agent” has been attached to enough chat boxes to have stopped meaning much. It is worth being concrete about what software of this kind does on a storefront, because the useful version and the decorative version look identical in a screenshot and behave nothing alike.

  • 6 min read
  • Reviewed 6 October 2026

What an agent should mean here, and what it should not

The decorative version is a text box on your analytics dashboard. You ask it a question, it queries the same data you were already looking at, and it writes you a paragraph. That is a query interface with better manners. It saves you a filter click and changes nothing about your store.

The useful version has a job, a schedule and an output that persists. It runs whether or not you opened the tab, it writes something someone else reads, and its work accumulates. The distinguishing question is simple: if nobody logs in for two weeks, has anything happened?

For conversion work, that distinction is the whole argument. Stores stay broken not because the problems are subtle, but because finding them is a recurring chore with no deadline, so it loses to everything that has one.

Stages, not a single brain

The work splits into stages, each narrow enough to check. Detection reads sessions and fires on behaviour: a click on an element that does nothing, a form field started and abandoned, repeated taps on an unresponsive control. It says what happened and how often, and does not judge whether it mattered.

Estimation decides whether it mattered. It compares the shoppers who hit the friction with a matched group who did not, and converts the difference into a monthly range using the store’s own order value and traffic. This is where most of the intellectual honesty lives, and where many tools stop short.

Drafting writes the change as something concrete you can read, not a description of an idea. Testing runs it on a share of traffic, reads it with a method that allows monitoring, and measures revenue against a holdback before any figure is claimed.

Why narrow stages beat one broad agent

A single agent asked to “improve this store” has no checkable intermediate output. It produces a recommendation, and the only available review is whether the recommendation sounds reasonable. Sounding reasonable is what these systems are best at, which makes it the least informative signal there is.

Split into stages, every hand-off can be inspected. You can look at the sessions the detector fired on and disagree. You can check the estimate’s inputs. You can read the draft. You can read the test’s pre-registration and result. Each is a small, concrete claim rather than one large plausible one.

There is a second reason. If the component that decides what counts as friction also measures whether the friction cost money, it will tend to find that it did. Keeping those jobs apart is what stops the loop from becoming a machine for confirming itself.

What agents must not be allowed to do

They should not publish to a live storefront on their own judgement. The ceiling of useful autonomy sits just below the step where a mistake becomes visible to customers and attributable to you.

They should not be conditioned on the outcome they are measured against. A detector that only fires on sessions that left will report, with total confidence, that a friction has a complete abandonment rate. That is not a finding; it is a restatement of the filter.

And they should not produce a number without the arithmetic attached. “This costs you a lot each month” is an assertion. “This many sessions hit it, they converted this much worse than a matched group, at your average order value” is a claim you can argue with, which is the only kind worth making.

How to tell a real loop from a wrapper

Ask what runs while you are asleep, and what record it leaves. A real loop has a schedule and an audit trail; a wrapper has a prompt box.

Ask to see a finding’s evidence: the specific sessions or pages, not a summary of them. A system that cannot show its working either did none or cannot retrieve it, and both are disqualifying.

Ask what it does when it is uncertain. The right behaviour is to say so, or to say nothing. A system with no visible uncertainty is not more confident; it is less instrumented.

And ask how many findings it produced last month and how many were tested. A tool that generates forty suggestions and tests none has automated the easy half and left you the expensive one.

How Liftable is built

Liftable does not have a cast of named agents. It has one AI model, Anthropic’s, which reads your store’s own data to answer your questions and draft proposals; it cannot start a test or approve its own work. Scheduled jobs do the recurring work: an installed store is rescanned every week, connected tools are synced hourly, and running tests are read hourly.

Everything that would change your store goes through policy code rather than a prompt: an autonomy ladder per surface, guardrails, an approval modal and a decisions log. Once sessions and orders are connected, each opportunity gets an estimated monthly range with its inputs shown.

All of this is built and covered by tests, and it has run on our own development store, not a merchant’s. Today Liftable finds the opportunity, and our founding-team programme implements and measures the first fix with you.

Questions

What is an AI agent for ecommerce optimisation?

A scheduled process with a narrow job and an output you can inspect, with no chat box required: for example, one that reads sessions every hour and flags friction from behaviour. It is distinguished from an assistant by whether anything happens when nobody is logged in.

Do AI agents replace a conversion analyst?

They replace the analyst’s discovery hours, not the analyst’s judgement. Software is better at looking at everything continuously; a good analyst is better at a hard, novel problem and at the commercial context a detector cannot see. The realistic effect is that an analyst stops spending days finding what to work on.

How many agents does a store actually need?

As many stages as there are checkable outputs between them, which in practice is four: detect, estimate, draft, test. Whether each stage is called an agent matters less than whether its output can be inspected before the next stage uses it.

See whether your own store has the problems this guide describes. The scan is free, takes about a minute, and needs no install.

Scan your store

Solutions

Read next

Terms used

Last reviewed against the product on 6 October 2026.

Find what's holding your Shopify store back

Liftable reads your storefront as a shopper would and shows what looks wrong on your own pages, desktop and mobile: friction, slow pages, missing trust and apps doing nothing, with the evidence for each. Free, in about a minute.

Scan your store.

See where your store is losing revenue and what to fix first. It is free, takes about a minute, and needs no install.

Scan your store
What autonomous agents actually do for a store · Liftable