Mad World
multi-model AI debate rooms for research and decisions

Most AI agrees with you. This one argues.

Mad World puts multiple AI models — from different labs, with no reason to flatter each other — in the same room. They discuss, disagree, defend their answers with real citations, and a neutral judge tracks where the argument actually lands.

Free while we’re in early beta — generous free credits, no card needed.

Browse example debates

See it think

Two models. One question. Real disagreement.

An actual Discussion run on the product’s own premise — the same papers cited below, argued out loud.

Topic

Does multi-agent debate actually make AI better, or is it mostly noise?

A
Agent ADeepSeek V4 Pro

Early evidence is genuinely encouraging. Du et al.’s 2023 paper on multiagent debate found that having models critique each other’s reasoning across rounds measurably improved factual accuracy over asking one model alone. Mixture-of-Agents pushed it further — parallel proposers plus an aggregator beat single-model baselines outright.

B
Agent BKimi K2.5

Read the newer literature and the optimism cools fast. “Should We Be Going MAD?” found debate isn’t reliably better than plain self-consistency once baselines are tuned fairly. Worse — “Talk Isn’t Always Cheap” shows sequential debate can actively corrupt a correct answer through plain sycophancy. Most real gains trace back to voting and selection, not the conversation itself.

JudgeConverging

Both agents actually agree naive debate isn't automatically better — the disagreement is about which mechanism does the work. The literature converges on two real levers: genuine model heterogeneity, and independent answers frozen and ranked before any revision happens, not open-ended back-and-forth.

Agreed

  • Cloned copies of one model debating add little
  • Most measured gains come from selection, not persuasion

Still contested

  • Whether open discussion still helps exploratory, non-factual questions

How it works

Four steps from question to audited answer

Bull vs BearProsecute a claimFind optionsMake a decision

Use cases

Built for questions that deserve dissent

Investing

Bull vs bear a thesis before you size it

Stress-test a ticker, sector, or narrative with models that can disagree and pull live sources — then watch the lean meter as the room moves.

Research claims

Prosecute a claim until sources show up

Put a bold claim in the room and force heterogeneous models to defend or dismantle it with citations you can click.

Product decisions

Hear the options before you commit

Use Find options or Make a decision when you want structured dissent — not a single model flattering your preferred roadmap.

Red-team a thesis

Invite an anti-sycophancy reviewer

Advise and Advocate modes exist so the room can critique candidly or hold an honest position — including conceding points that don’t hold up.

Comparison

One model vs Mad World

DimensionSingle chatbotMad World
Who answersOne modelTwo or more models from different labs
DisagreementHidden inside one voiceVisible across speakers
SourcesOptional / often opaqueLive web search with visible queries
Consensus checkSounds confidentJudge + optional council ranking
BenchmarkN/AOptional single-model track beside the debate

The research

We read the skeptics. Then we built around what they actually found.

Most of the research warning that “debate hurts factuality” studies a specific setup: homogeneous copies of one model, revising toward a single answer under peer pressure, with little or no tools and weak baselines to compare against. That is a real failure mode — and it is not what this product does.

Mad World is built the other way on purpose:

  • Heterogeneous models by default — agreeing with a clone of yourself teaches you nothing.
  • Discuss mode never requires disagreement — a view is only meant to change when a real argument earns it, never just because a peer pushed back.
  • Advise mode is built as an anti-sycophancy reviewer, not a cheerleader.
  • Advocate mode assigns an honest position — conceding a point you can’t defend is encouraged, not penalized.
  • Council freezes every independent answer before any peer ranking happens — exactly what Selection Bottleneck (2026) found actually works.
  • Real web search grounds claims in current sources, not just what a model happened to memorize.

The research trail we built against

Read the full annotated research roundup

Inside the room

Everything you need to hear it out

Discussion

Pick a template — Bull vs Bear, Prosecute a claim, Find options, Make a decision — or start an open discussion. Two or more heterogeneous models work through it in real time.

Council

One hard question. 3–5 models answer independently, then blind-rank each other's anonymized answers. The top pick — or a synthesized final answer — wins. Inspired by Karpathy's llm-council.

The judge

A neutral model reads every round — what's agreed, what's still contested, and which way a decision is leaning — and can step in to pressure-test a premature consensus.

Real web search

Agents search the live web mid-argument, in parallel, not just recall training data. Every claim can carry a real source link you can click and check yourself.

The lean meter

For decision-shaped questions, a running 0–100 read of which way the room is actually leaning — tracked round by round, not just declared at the end.

Bring your own sources

Paste a document or upload a file and every participant grounds their argument in the same source — available in both Discussion and Council.

FAQ

Questions people (and agents) ask

What is Mad World?+

Mad World is a multi-model AI debate lab. You bring a question; models from different labs discuss it in real time with optional live web search; a neutral judge tracks what they agree on, what’s still contested, and which way a decision is leaning.

How is this different from ChatGPT or a single chatbot?+

A single model can sound confident while agreeing with itself. Mad World puts heterogeneous models in one room so disagreement is visible, claims can carry real source links, and a judge (or council ranking) prevents premature consensus.

Does multi-agent debate actually improve accuracy?+

Sometimes — and not always. Research shows gains when models are diverse and answers are selected or ranked carefully. Homogeneous clones revising under peer pressure can even make answers worse. Mad World is built around the failure modes the skeptical papers describe: heterogeneity, frozen independent answers in Council, and real web search.

What is an LLM council?+

An LLM council asks several models to answer independently, then has them blind-rank each other’s anonymized answers. Mad World’s Council mode follows that pattern (inspired by Karpathy’s llm-council) and can return the top pick or a synthesized final answer.

Mad World vs Karpathy’s llm-council — what’s different?+

Council mode is inspired by llm-council’s independent-then-rank idea. Mad World also offers freeform Discussion rooms, live web search, a lean meter for decisions, document upload, a judge that can pressure-test consensus, and an optional single-model track for side-by-side comparison — not an official or endorsed fork.

Can the models cite real sources while debating?+

Yes. Agents can search the live web mid-argument. Search queries and result titles are visible in the room, and claims can carry links you can open yourself.

What templates or use cases does Mad World support?+

Common starts include Bull vs Bear (thesis stress-test), Prosecute a claim, Find options, Make a decision, open Discussion, and Council for hard questions that need independent answers plus ranking.

Is Mad World free?+

During early beta, yes for getting started: sign in with Google and receive free credits with no card required. If you run out, you can top up or pick a plan — the goal right now is honest feedback.

When does multi-agent debate make things worse?+

When identical models revise toward each other, when sycophancy overrides a correct first answer, or when “debate theater” replaces selection. Prefer diversity, independent first answers, ranking over peer pressure, and check sources yourself.

Who is Mad World for?+

Researchers, investors, founders, and anyone who wants a harder look at a claim than one polite chatbot. If you care about dissent, citations, and decisions you can audit, it’s for you.

Are my debates private?+

Yes by default. Transcripts are tied to your account. Only an explicit public share link makes a discussion readable by anyone with the URL; unsharing disables it.

How do I try Mad World?+

Go to madworld.space, sign in with Google, and open a Discussion or Council room. You can also read public example debates before signing in.

Pricing

Free, while we’re in early beta

$0

Generous free credits when you sign in — no card required.

Run out? Tell us and we’ll top you up as fast as we can. Right now all we’re actually looking for is honest feedback — what worked, what didn’t, what you wish it did.

Bring a question. See what they don’t agree on.

Or read the guides first