Delphi Pythia Take the survey →

A computational mini-public

The better argument, measured.

Delphi Pythia gathers a public to argue a question, then measures two things: which arguments people rate as persuasive, and which ones change minds.

The instruments

Two instruments, one discipline

Choose an instrument

The problem

What the existing tools miss

Surveys measure what people believe. They are poor at explaining why anyone changed their mind. Deliberative forums surface real arguments, but they do not scale: a citizens' assembly is slow, costly, and small. Online discussion produces a great deal of text and very little measurement.

The same gap runs under all three. What people say and rate is an elicited signal, and it tells you which arguments sound convincing when someone is asked to judge them. What changes a position is a caused effect. The two come apart, and most instruments only ever see the first.

The solution

Discover, then verify

Delphi Pythia runs the deliberation and the experiment in the same instrument. On any contested question it works in two passes. Wave one discovers the arguments a public is raising, the positions they cluster into, and which of those positions is gaining ground as people rate one another's points. Wave two freezes that argument map and runs a randomized experiment on it, which shows which individual arguments are doing the moving.

That is what makes real deliberation practical at a scale citizens' assemblies cannot reach, without losing rigour. Separating each argument from its author means status and authority count for less. The analysis gives trustworthy estimates of how opinion moved, and it brings out the rare arguments that cross divides and win people over. It is built as a consultative instrument, a way to listen to a public more carefully.

How it measures

How the measurement works

Delphi Pythia is a computational mini-public, built on a long tradition of structured public deliberation: the Athenian assembly, Madison's call to refine and enlarge public opinion, Habermas's ideal of the unforced force of the better argument, and the anonymous, iterative rounds of the Delphi method. That lineage is why the instrument works the way it does. People contribute their own arguments, then read and rate everyone else's.

The first wave is open deliberation. People write arguments in their own words, and the system clusters them into an argument map as they arrive. An adaptive selector (exploration sampling, a Thompson-sampling variant) chooses which argument you see next, concentrating on the ones whose standing is still unresolved, while a built-in floor keeps every argument in rotation so none gets buried early. That tells us what people rate as persuasive, which is not always what moves them.

The second wave takes the same argument map, now frozen, and shows people randomized combinations of arguments, the way a lab experiment would. Position is measured before and after, so the shift can be read directly instead of inferred from a rating. Significance comes from reshuffling the real data thousands of times, which assumes nothing about the shape of the distribution, and the ranking reports its own uncertainty.

When to use

When the question has sides

Delphi Pythia is the instrument for a question people already argue about.

1,000–2,000 respondents, across all camps
  • People already hold a position. The question has sides, and respondents can say which one they are on: a policy, a product direction, a referendum.
  • You need to know what moved people. Ratings tell you how convincing an argument sounded to someone asked to judge it. The second wave measures the position itself, before and after.
  • You care how it lands with the other side. Whether an argument crosses a divide or hardens the camp against it only shows up when positions are declared and kept apart.

The problem

What research actually costs

Good market research is expensive in ways that rarely show up on the invoice. Recruiting the right respondents is hard, and the people you most need to hear from are the busiest ones. Every extra wave adds coordination, vendor management and weeks of waiting. So teams run one big study, get back a stack of ratings where everything looks important, and still have to guess what to lead with. The knowledge they needed was in their market the whole time. The expensive part is getting it out in a usable order.

The solution

Map, then prioritize

Delphi Compact is built for that economics. In the discovery wave respondents answer in their own words, and each session rates what earlier sessions raised. An adaptive allocator decides where the next rating is most useful, so a study of one hundred and fifty people does the work that usually takes several hundred. You get the arguments of your market from the bottom up, not from a workshop guess. The measurement wave then takes the short list that survived and runs structured trade-off choices on it. The result is not another wall of scores. It is a ranked list of what matters, ready for a positioning decision.

How it measures

The same measurement discipline

The efficiency does not come at the price of rigour. The allocator records the odds it gave every item, and the estimates correct for those odds, so adaptive routing cannot quietly bias the map. Trade-off choices are randomized, which is what lets a small sample produce a defensible ranking. Ranks come as a set the data can support rather than a single confident order. When the study cannot answer something, the report says so instead of rounding up.

When to use

When nobody has a side yet

Delphi Compact is the instrument for a market that has not been asked yet.

≈300 completes, end to end across both waves
  • Preferences are still forming. People work out what they want as they answer, which is why the instrument asks before it shows them anything.
  • You need an order you can act on. The measurement wave returns a short ranked list with its own uncertainty attached.
  • Recruitment is the constraint. Three hundred completes carry a full study, because the adaptive allocator spends each rating where it is worth most.

At a glance

Side by side

Pythia Compact
Question Which arguments move a formed position? Which criteria matter, and how much?
Positions Declared stances, partitioned None assumed — structure is derived
Waves Discover → randomized verification Discover → randomized trade-offs
Sample ≈1,000–2,000 total ≈300 total (150–300 per wave)
Claim Persuasion, causally identified Salience and priority; never persuasion

The demo case

Where the debate is heading

The instrument is running end to end on a demonstration case: how AI should be governed. See the live results →

These findings come from our first human pilot: 112 people worked through the question of how AI should be governed. About one in six (17%) changed their stated position along the way, and average confidence in the position people landed on rose by about half a point on the five-point scale. A few patterns are already clear.

BEFORE AFTER 45% 53% International body +8 40% 37% National gov't -3 15% 11% Industry self-rule -4

Initial vs. final position, first human pilot · n = 112

Over the course of this pilot, support moved toward an international body to coordinate AI rules. Of the three positions on the table, it gained the most ground.

  • Accountability is what people rate highest.

    The argument rated most persuasive across the divide was that companies should not be allowed to run AI without legally binding government oversight. Even people who arrived favouring a different approach rated it 3.9 / 5. The weakest was the call to let industry regulate itself.

  • There is real common ground.

    Two arguments are rated well by every camp: that a single international body should set common rules all countries follow, and that companies should face clear standards, public transparency, independent audits, and real liability. Either could anchor a position most people could live with.

  • Rated highest is not the same as most effective.

    This pilot measured what most polling and message-testing measures: which arguments people rate highly after reading them. It does not tell us which of those arguments moved someone's position, as opposed to simply being on screen when the position moved. That is what the second wave is for: a randomized experiment on the same argument map. See it running now, above.

The same two-wave process can run on any contested question, from a policy to a product to a referendum: discover the arguments a public is making, see which way opinion is moving, and verify which specific arguments are shifting positions. Ratings alone cannot do the last part.

The cases

The same instrument, different questions

Each one runs the same two waves on a different contested question. Pick one to see how its argument map moved, or take the study yourself.

Take one of the studies yourself. You state a position, weigh the strongest arguments from every side, and watch where you land as the room moves.

or see the live results 📊