Investing
Bull vs bear a thesis before you size it
Stress-test a ticker, sector, or narrative with models that can disagree and pull live sources — then watch the lean meter as the room moves.
Mad World puts multiple AI models — from different labs, with no reason to flatter each other — in the same room. They discuss, disagree, defend their answers with real citations, and a neutral judge tracks where the argument actually lands.
Free while we’re in early beta — generous free credits, no card needed.
Browse example debatesSee it think
An actual Discussion run on the product’s own premise — the same papers cited below, argued out loud.
Topic
Early evidence is genuinely encouraging. Du et al.’s 2023 paper on multiagent debate found that having models critique each other’s reasoning across rounds measurably improved factual accuracy over asking one model alone. Mixture-of-Agents pushed it further — parallel proposers plus an aggregator beat single-model baselines outright.
Read the newer literature and the optimism cools fast. “Should We Be Going MAD?” found debate isn’t reliably better than plain self-consistency once baselines are tuned fairly. Worse — “Talk Isn’t Always Cheap” shows sequential debate can actively corrupt a correct answer through plain sycophancy. Most real gains trace back to voting and selection, not the conversation itself.
JudgeConverging
Both agents actually agree naive debate isn't automatically better — the disagreement is about which mechanism does the work. The literature converges on two real levers: genuine model heterogeneity, and independent answers frozen and ranked before any revision happens, not open-ended back-and-forth.
Agreed
Still contested
Topic
15
Strongly Don't pursue
Your core assumption has a crack in it. You’re betting that a cheaper, less capable model can do something the main coding agent — which actually wrote the code and has full context — supposedly can’t: explain what just happened clearly. The difficult part isn’t the visualization layer — it’s the semantic extraction, and that’s a reasoning problem, not a rendering problem.
Advisor A is right that semantic extraction is the hard part, but you’re facing an even harsher reality: you’re arbitraging a UX gap the platforms will close themselves. Cursor and Windsurf already own the agent’s intent and reasoning traces; you’re trying to reverse-engineer meaning from raw diffs. The moment an incumbent ships a native “what changed” view, your sidecar becomes redundant.
JudgeConverging
Both advisors agree the sidecar is structurally flawed but diagnose different fatal flaws: decision-detection from diffs is effectively AI-complete, and incumbents will natively obsolete the tool with structured reasoning traces of their own.
Agreed
Still contested
Topic
58
Leaning Bull
The Q1 FY27 results released just a week ago reinforce HDFC Bank’s structural strengths: deposit growth accelerated to 13.3% YoY, well ahead of typical Q1 seasonality, and management confirmed the bank is on track for a 40–50 bps reduction in its cost of funds. That deposit momentum is the hardest thing for competitors to replicate.
The 5% profit growth you’re celebrating is actually a significant miss that exposes HDFC Bank’s structural deceleration post-merger — this is not “modest optical” weakness but a fundamental reset. NII grew just 6.7% and NIM remained compressed, confirming the bank can’t protect yields in a competitive lending environment.
JudgeConverging
The debate has clarified key factual disputes, with the Bear conceding that Q1 deposit growth outpaced the system. The discussion now centers on interpretation: whether strong liability momentum will translate into promised margin expansion, or whether structural post-merger dilution permanently impairs profitability.
Agreed
Still contested
How it works
Use cases
Investing
Stress-test a ticker, sector, or narrative with models that can disagree and pull live sources — then watch the lean meter as the room moves.
Research claims
Put a bold claim in the room and force heterogeneous models to defend or dismantle it with citations you can click.
Product decisions
Use Find options or Make a decision when you want structured dissent — not a single model flattering your preferred roadmap.
Red-team a thesis
Advise and Advocate modes exist so the room can critique candidly or hold an honest position — including conceding points that don’t hold up.
Comparison
| Dimension | Single chatbot | Mad World |
|---|---|---|
| Who answers | One model | Two or more models from different labs |
| Disagreement | Hidden inside one voice | Visible across speakers |
| Sources | Optional / often opaque | Live web search with visible queries |
| Consensus check | Sounds confident | Judge + optional council ranking |
| Benchmark | N/A | Optional single-model track beside the debate |
The research
Most of the research warning that “debate hurts factuality” studies a specific setup: homogeneous copies of one model, revising toward a single answer under peer pressure, with little or no tools and weak baselines to compare against. That is a real failure mode — and it is not what this product does.
Mad World is built the other way on purpose:
The research trail we built against
Early evidence that critique across rounds can beat asking one model alone on factual tasks.
Diversity of models plus confidence-weighted voting outperforms clones of one model.
Parallel proposers with an aggregator beat single-model baselines on several benchmarks.
Debate is not automatically better than strong self-consistency — baselines matter.
A skeptical read: many MAD gains shrink once evaluation is tighter and fairer.
Sequential peer pressure can talk a correct model into a wrong answer via sycophancy.
Freezing independent answers and ranking them often beats open-ended synthesis.
Inside the room
Discussion
Pick a template — Bull vs Bear, Prosecute a claim, Find options, Make a decision — or start an open discussion. Two or more heterogeneous models work through it in real time.
Council
One hard question. 3–5 models answer independently, then blind-rank each other's anonymized answers. The top pick — or a synthesized final answer — wins. Inspired by Karpathy's llm-council.
The judge
A neutral model reads every round — what's agreed, what's still contested, and which way a decision is leaning — and can step in to pressure-test a premature consensus.
Real web search
Agents search the live web mid-argument, in parallel, not just recall training data. Every claim can carry a real source link you can click and check yourself.
The lean meter
For decision-shaped questions, a running 0–100 read of which way the room is actually leaning — tracked round by round, not just declared at the end.
Bring your own sources
Paste a document or upload a file and every participant grounds their argument in the same source — available in both Discussion and Council.
Public transcripts
Curated public shares we own (or have permission to promote) — primary sources AI engines and humans can cite.
The product’s own premise argued out loud — when multi-agent rooms help, and when they don’t.
Exploring an AI coding / visualization product idea through multi-model discussion.
Do expert-style prompts still matter when multiple models are in the room?
FAQ
Mad World is a multi-model AI debate lab. You bring a question; models from different labs discuss it in real time with optional live web search; a neutral judge tracks what they agree on, what’s still contested, and which way a decision is leaning.
A single model can sound confident while agreeing with itself. Mad World puts heterogeneous models in one room so disagreement is visible, claims can carry real source links, and a judge (or council ranking) prevents premature consensus.
Sometimes — and not always. Research shows gains when models are diverse and answers are selected or ranked carefully. Homogeneous clones revising under peer pressure can even make answers worse. Mad World is built around the failure modes the skeptical papers describe: heterogeneity, frozen independent answers in Council, and real web search.
An LLM council asks several models to answer independently, then has them blind-rank each other’s anonymized answers. Mad World’s Council mode follows that pattern (inspired by Karpathy’s llm-council) and can return the top pick or a synthesized final answer.
Council mode is inspired by llm-council’s independent-then-rank idea. Mad World also offers freeform Discussion rooms, live web search, a lean meter for decisions, document upload, a judge that can pressure-test consensus, and an optional single-model track for side-by-side comparison — not an official or endorsed fork.
Yes. Agents can search the live web mid-argument. Search queries and result titles are visible in the room, and claims can carry links you can open yourself.
Common starts include Bull vs Bear (thesis stress-test), Prosecute a claim, Find options, Make a decision, open Discussion, and Council for hard questions that need independent answers plus ranking.
During early beta, yes for getting started: sign in with Google and receive free credits with no card required. If you run out, you can top up or pick a plan — the goal right now is honest feedback.
When identical models revise toward each other, when sycophancy overrides a correct first answer, or when “debate theater” replaces selection. Prefer diversity, independent first answers, ranking over peer pressure, and check sources yourself.
Researchers, investors, founders, and anyone who wants a harder look at a claim than one polite chatbot. If you care about dissent, citations, and decisions you can audit, it’s for you.
Yes by default. Transcripts are tied to your account. Only an explicit public share link makes a discussion readable by anyone with the URL; unsharing disables it.
Go to madworld.space, sign in with Google, and open a Discussion or Council room. You can also read public example debates before signing in.
Pricing
$0
Generous free credits when you sign in — no card required.
Run out? Tell us and we’ll top you up as fast as we can. Right now all we’re actually looking for is honest feedback — what worked, what didn’t, what you wish it did.