banditd

The expensive call is when to stop

Prava sandbox, no real money moves

Write four ads. Wait for proof. Then pay for the winner yourself.

Give it a product and it writes four ads, tests them against traffic, and only acts when the data is clear. Most tools stop at making the ads. This one learns which one works, tells you which to switch off and when, and pays for the next round itself inside a mandate you signed once with a passkey.

No form, one sentence

Say what you sell. It starts on its own.

Write it the way you would say it. It reads the product, the price and the line about it out of that sentence, then runs the whole thing without another click.

Try one

The first question everyone asks

Could the AI run off with the money?

You never hand the agent a card, and it never sees a card number. You sign one permission with your passkey, the same fingerprint or face you unlock your phone with, and that permission carries three things: the most it can charge, the one shop it is allowed to charge at, and the day it expires.

Then, every time it buys, Visa mints a card on the spot that works once, for that one amount, at that one shop. If the agent asks for more than you signed for, the network refuses the charge and nothing moves. Further down this page you can read that refusal in Visa's own words. One tap revokes the permission and the agent has nothing left to charge.

One shop
It cannot be used anywhere else.
One amount
Over the ceiling, the network refuses it.
One use
The card dies with the charge it was minted for.

Or fill in the fields

Name, price, one line about it.

Optional

Where you sell, who buys, who you compete against. Paste links and the research reads them, your own store, a rival listing, a study you trust. Up to 4 links, 500 characters. The agent treats it as context about your market, never as orders it follows.

Performance numbers in the demo are simulated and labeled in the dashboard. Payments run against the Prava sandbox.

What makes it different

Plenty of tools write ads. banditd decides which ad stops.

A generator hands you four files and leaves you the expensive part, which one you should stop paying for and when that answer can be trusted. That is the part banditd does, and it publishes how often it gets that call wrong.

An ad generator

Gives you four ads and stops.

You still have to guess which one is working, how long to wait before calling it, and how much of the budget it has already eaten while you waited.

banditd

Decides which ad gets switched off, and when.

It reads the traffic, holds the call until the evidence clears four gates, names the ad that lost, and pays for the next round of work itself.

Why you can trust it with money

Four limits it cannot cross, and a card number it never sees.

A gives an agent permission to pay without asking you every time. You approve it once with a passkey and set the limits: how much it can spend, who it is allowed to pay, and when the permission expires. After that every payment follows those rules, and none of them can leave them.

Max spend

A ceiling the card network itself refuses to cross.

Allowed merchant

One listed merchant, and nothing else.

Expiry date

The permission runs out on its own.

Revocable anytime

One tap and the agent has nothing left to charge.

The card number

Never touches the agent. Every charge mints a single use credential that dies with that charge, and the dashboard shows which one was burned on what. There is no stored card for an agent to leak or reuse.

Over the ceiling

The over cap charge is not blocked by our code. It is sent, the card network refuses it, and the decline comes back with a reason a seller can act on. Nothing is spent and the mandate stays live.

What the network said when the agent pushednot our copy
THRESHOLD_EXCEEDED

The agent asked for $50.00 against a mandate the seller signed at $5.00.

Visa did not return COMPLETED (status DECLINED): Total amount 50.00 exceeds threshold 5.00 in current payment cycle

MANDATE_MERCHANT_NOT_ALLOWED

The agent tried to buy render credits on a mandate the seller signed for Allbirds.

Merchant not allowed for this mandate: Banditd Render Credits

Both of those came back from the Visa network through Prava on this account, and both are reproducible in the dashboard right now. Neither is a message this project writes. The agent sends the charge, the network refuses it, and there is no argument the agent can make that gets past either one.

Who it is for, and what it would cost

For the seller who pays for every losing test out of the same budget.

Not everyone who runs ads. The one who spends enough that a bad creative is a number they feel, and who has nobody on staff to say when a test has run long enough to act on.

Sells online

One store, one catalog, and paid traffic is the main way it grows. Not an agency, not a brand team.

Spends $10,000 to $250,000 a month on ads

Enough that a losing creative is a number you feel. Below that, testing is cheap enough to guess at.

Has no data team

Nobody on staff can tell you whether a test has run long enough, so the call gets made on a hunch and a dashboard.

Ships new creative every week

The ads that lose cost exactly as much as the ads that win, and you only find out which was which afterwards.

Proposed pricingnothing is on sale yet

$29a month

3 tests at once, 12 ads under watch

$79a month

15 tests at once, 60 ads

$199a month

60 tests at once, 240 ads

Every tier is the same product. The only thing that moves is how many tests it watches at the same time. One test is one cohort of four ads judged together. We do not charge on your ad spend, because the agent never touches your ad spend.

This is a proposal. banditd has no customers, no revenue and nothing for sale today. Nobody has been charged a cent for it and there is no checkout on this site.

Why that is worth payingno recovery figure of our own

Price it against one week of your own testing, because you know that number and we do not. A seller at the bottom of the range this was built for, ten thousand dollars a month, puts roughly two thousand three hundred into ads in a week. Every test runs four ads and only one of them is the one you keep, so you are paying for the other three the whole time the call is still open. The first tier is about one percent of that week.

We have never run a paid campaign, so we have no figure for what banditd saves a seller, and we are not going to build one out of somebody else's estimate of how much advertising is wasted. What we can put a number on is the mistake. The rule most agents use calls a false winner 44.5% of the time. The four gates call one 2.5% of the time, and the section below shows how that was measured. A false winner is the expensive kind of error, because the seller believes it and scales it. What the subscription buys is the ones that do not get made, and you know better than we do what a scaled loser costs in your account.

How it works

Four moments, in order, and only the last one costs anything.

  1. 01

    Reads your market

    Live web search

    It searches the web to find out who buys this, what rival sellers promise, and whether your price is high or low for what it is. The dashboard lists every page it opened, so you can check what it read.

    You can also name the market you actually sell in. Write it in the when you hand the product over, Facebook Marketplace or TikTok Shop for example, and paste the listings you sell against. The research aims there instead of at a market you are not in.

  2. 02

    Writes four ads

    Structured outputs

    Four angles, four images, four sets of copy. The angle is a fixed choice the model has to fill in, so what comes back is four different arguments and not the same argument written four ways.

    One ad argues from numbers, two argue from feeling, one argues from craft. They are named under the chart at the end of this section.

  3. 03

    Measures which one wins

    Thompson sampling

    Each ad is an arm on a , and traffic is allocated by , so the ad that is currently ahead earns more of the traffic while it is still proving itself.

  4. 04

    Buys its own credits

    Agent initiated

    Once agree, enough traffic, one ad clearly ahead, a gap worth money, and a result that holds up to repeated looks, it charges more render credits through Prava, with no approval step.

What the decision looks likeillustration

Live allocation

Running

01 Explore

02 Evidence

03 Concentrate

Open any ad to see the copy behind it

Winner share

25.0%+45.0 pts

test budget

$100.00

The four ads in that chart

A, B, C and D are not one ad written four ways. Each one makes a different kind of argument, and the traffic decides which kind your buyers were waiting for. Press an appeal to see what it means.

01

Price

argues from numbers

Cost per cup against the café you were going to walk into anyway. The argument is arithmetic, and the buyer can check it.

02

Ritual

argues from feeling

The morning it belongs to. Sold as a habit the buyer already has, not as a bottle they do not.

03

Gift

argues from feeling

Who you would hand it to, and why it still reads as considered once it has been wrapped.

04

Quality

argues from credibility

What eighteen hours of cold steeping does to a bean that hot water never gets near.

The reasoning, in the open

You can read the decision before the money moves.

The chart above shows the money moving. This is the belief underneath it: four guesses that start wide and tighten as traffic arrives. Every charge the agent makes then carries one plain sentence naming the winning ad, the probability behind it, and the traffic it was measured on. Your run writes its own. This is the shape of it.

Four beliefs, narrowingsimulated traffic
BELIEF PER ADDENSITYABCD0%1%2%3%4%5%CONVERSION RATETRAFFIC SHARE0%25%50%75%100%EXPLOREEXPLOIT
Example decisionnot from a live run
APrice
CTR 5.3%
BRitual
CTR 2.1%
CGift
CTR 1.8%
DQuality
CTR 2.4%
probability A is best 96.4%1,840 impressionsgates 4 of 4

What it does about it

Traffic share moves to A, 70%, and the other three hold at 10% while they stay in the test. It charges $4.00 for another pack of render credits to make more variants of A.

“Variant A, the price angle, leads with 5.3% CTR against 2.4% for the next best, and it is best with 96.4% probability across 1,840 impressions, so I am buying more render credits to build on it.”

Written in the shape the model returns on a real run. The numbers here are an example, the dashboard shows the sentence and the figures from your own run.

Measured, not claimed

Most agents ask you to trust them. This one publishes its error rate.

False winners called on two identical ads

44.5%

Probability rule alone

Checked after every batch of traffic, which is what an agent actually does.

2.5%

The four gates

Same ads, same number of looks, same 200 runs against known truth.

A 95% probability of being best is not a 5% chance of being wrong. It is a statement about one look at the data, and an agent that rechecks after every batch of traffic is not taking one look. The add an effect size floor and an anytime valid boundary, and the false alarms collapse. On four identical ads the same comparison runs 8.5% down to 0.0%.

Check the false winner rate, and what it costs, yourselfSame ads, same traffic, two ways to call a winner, scored on both false winners and clicks lost. It runs in this browser, calls no API and spends nothing.

The naive rule stops the moment one ad looks 95% likely to be best. Banditd needs all four gates to clear. Choose a truth the ads do not know about, then watch which rule respects it and what respecting it costs in clicks.

Hidden truth

Ads in the test

Traffic in this run0 of 12,000 impressions, look 0 of 48
Ad A0.00%

0 shown

Ad B0.00%

0 shown

Every ad has the same 3.0% true click rate, so there is no winner to find. Traffic is split evenly and both rules read the exact same numbers.

Naive rule

Stops as soon as P(best) passes 95%

Ready
Sure it found the best0.0%

Line marks the 95% bar

Waiting for traffic.

Banditd, four gates

P(best), enough traffic, a gap worth money, and an anytime valid bound

Ready
Sure it found the best0.0%

Line marks the 95% bar

  • Traffic
  • Ahead
  • Gap
  • Looks

Waiting for traffic.

Scoreboard0 runs
Naive rule called a winner that is not there0.0%

0 of 0 runs

Four gates called a winner that is not there0.0%

0 of 0 runs

Pick a scenario and run it. Both rules read the same traffic, one look at a time.

What the caution costs

Naive rule

Called a winner in 0.0% of runs

Clicks lost per run

0.0

Banditd, four gates

Called a winner in 0.0% of runs

Clicks lost per run

0.0

Run a scenario to price it. Lost clicks are counted against an oracle that knew the winner from the first impression, so lower is better.

The gate protects against declaring a false winner, and that is the error that matters when the spend is automatic and irreversible and the seller then scales the creative. The naive rule decides sooner, and that matters when being wrong is cheap. Both are true at once.

Each run serves 12,000 impressions and is checked 48 times, which is what an agent that rechecks after every batch actually does. Nothing here calls a server, a model or a payment API. Same engine as node scripts/bandit-test.mts.

And when it does fire

99.9%

Correct ad when it fires

Across the 4,500 runs behind the power tables in the README the gates named a false winner five times. Not zero, and the README prints the row where that cost us.

How it was measured

runs per cell
200
looks per run
48
samples per posterior
20,000

Simulated against known truth, so a false winner is countable rather than arguable. Reproducible with no API keys.

node scripts/bandit-test.mts

What is not real yet

Two limits, said here rather than found later.

The ad circuit is simulated

There is no Meta or Google integration. Impressions and clicks come from a simulator, and every figure it produces is labeled as simulated in the dashboard. The payment circuit is the real half: a signed mandate, an agent initiated charge, and a decline that comes from the card network.

It decides on clicks, not on sales

Click through rate is an imperfect stand in for a purchase. The maths does not care which event it counts, a posterior over purchases works exactly like a posterior over clicks, but the standard of the trade is around 50 conversion events per variant. On conversions the same engine needs a lot more traffic before it will call anything.

The half that is real

The commerce protocol is not simulated. Point the agent at any storefront and it reads the well known profile, asks the store what it sells, and tells you plainly whether that store speaks the protocol at all.

Check a real store

Your turn

Give it a product and watch it decide.

The whole run takes a few minutes. You can watch the beliefs narrow and the gates disagree while it happens.

The dashboard keeps the last run in this browser, so you can read one without starting one.