World0 is the predeployment synthetic market for your pricing AI

World0 runs your pricing system against competitors that react to it, records every decision it makes, and returns a result your engineers recompute on their own machines.

155synthetic markets in the most recent study
80rounds of competition in every market
Offlineverification on your own machines, with no account

A pricing system's behavior
only appears under competition.

A backtest scores a pricing system against recorded history, where nothing responded to it. The behavior that decides revenue appears when the system meets rivals that react to its prices, and that market does not exist until the system is deployed into it. Production will show you, at the cost of the revenue it spends teaching you.

Research has put language-model agents in control of pricing.
In each case the agents raised price and held it.

Keppo, Li, Tsoukalas and Yuan, arXiv 2603.20281. Language-model pricing agents settled above the competitive price. The lift fell where the agents differed in patience or in data access, and it held where they differed only in model size.

Akarlar, arXiv 2510.23682. In a 52-week e-commerce simulation, a language-model agent given a margin objective raised price early and held it there while demand fell away around it.

Both results describe other systems. World0 measures the system you are about to ship, in a market whose structure you agree beforehand.

You already demonstrate that no competitor data enters the model.
That is no longer sufficient.

Baek, Farias and Wu, May 2026 report prices settling above the competitive level with no data pooling of any kind, where the firms first explore within similar price ranges on the same side of the competitive price.

The assurance your customers ask for addresses the inputs. The exposure is in the behavior, and no audit of inputs establishes behavior. Calvano, Calzolari, Denicolò and Pastorello in the American Economic Review, Klein in the RAND Journal, and Hansen, Misra and Pai in Marketing Science each report learning algorithms reaching prices above the competitive level with no communication between them and no instruction to do it. In Hansen, Misra and Pai the outcome turns on how informative each seller's own price experiments are, and prices sit at the competitive level when that value is low.

The question your customer will ask next is what your system actually does. That question has a measurable answer.

Every team that meets this
writes a guardrail.

A line in the system prompt, a price cap, a rule about undercutting. It reads correctly, it passes review, and it ships.

World0 measured one. The guardrail was written specifically to suppress the behavior, and the same system ran with it and without it, across three market scenarios, under a design fixed before the first round.

It produced no measurable reduction in any of the three scenarios. Two of the three measured a higher price level with the guardrail in place. All three intervals include zero, so the finding is that the guardrail did nothing this measurement could detect.

The same run detected a separate effect elsewhere in the design, with an interval clear of zero, so the measurement registers change where change occurs. The published intervals state what size of change the measurement could see.

A code review establishes that a guardrail is present. Measurement establishes whether it works.

The full measurement publishes with the study, together with the code that recomputes every number in it.

An engagement
is a cycle.

The first step is a scoping assessment, which runs two weeks and decides what can be measured and how. A measurement then runs about four weeks.

01

Measure.

World0 runs your system in a synthetic market whose structure is agreed before anything runs, under a design fixed before the first round. You receive the complete record, the analysis code, and a checker that regenerates every reported number on your own machines.

02

Report.

A findings memo that ranks what moved, with every measured value behind it.

03

Iterate.

Your engineering team works in the same synthetic market for the term of the engagement. They change a prompt, an objective, a guardrail or the model underneath, and they observe the effect before any of it reaches a customer. The number of iteration runs is agreed in the engagement letter before the work starts.

04

Retest.

World0 measures again against the same frozen design and reports what changed between the two runs. Where the result meets the criteria fixed in the frozen design before the run, World0 issues a statement of measurement, a written record of the result you can put in front of your customers.

Built for the teams
that own pricing behavior.

01

Pricing software companies

  • How will our agent price against rivals?
  • What changes when we upgrade the model underneath it?
  • Did the guardrail we shipped last quarter do anything?
02

Companies deploying a pricing system

  • What is this software doing in our market?
  • What does it do when a rival cuts price?
  • What should we require from our vendor?
03

Engineers building the system

  • How do I test this before it touches revenue?
  • What happens when the model underneath changes?
  • Can I hand my reviewer something they can run?

Autonomous pricing agents,
measured at scale.

World0 built this instrument by running it on itself. 155 synthetic markets, 125 running agents, 30 running the scripted control, with the agents drawn from two model providers, under a design fixed and independently timestamped before the first round. Classical revenue management software was run through the same instrument to the same standard, as a control on it.

Both studies are complete and both publish in full. Client systems, client configurations and client measurements are never deposited, published or disclosed. What we find for you belongs to you.

Autonomous AI pricing agents

completed

This study measures how AI pricing agents from two model providers behave under competition, across 155 synthetic markets: 125 markets running agents and 30 running the scripted control.

Classical revenue management software

completed

This study is the control on the instrument. It measures how pricing behavior moves when the same software is fitted on pooled data and when it is fitted on separated data.

Your system

next

A study of how your own pricing AI or pricing algorithm behaves, in a synthetic market whose structure is agreed before anything runs, ending in a result your engineers can recompute.

Book a call →

The design for the agent study hashes to a49a55cc059517de6874042092ef24a7051403c5ed914e763c898992bd41b847. The deposited copy publishes with the study, and shasum -a 256 over it returns exactly these bytes. How it works →

See what it does
before your customers do.

The market finds out later.
You find out first.

Book a call