Research roundup: multi-agent debate
Key papers on multi-agent debate range from early factuality gains (Du et al.) to warnings that peer pressure can corrupt correct answers — Mad World designs around both.
Key papers on multi-agent debate range from early factuality gains (Du et al.) to warnings that peer pressure can corrupt correct answers — Mad World designs around both.
Annotated trail
- 2023 — Du et al. — Improving Factuality via Multiagent Debate — Early evidence that critique across rounds can beat asking one model alone on factual tasks.
- 2023 — ReConcile — heterogeneous agents, confidence-weighted voting — Diversity of models plus confidence-weighted voting outperforms clones of one model.
- 2024 — Mixture-of-Agents — parallel proposers + aggregator — Parallel proposers with an aggregator beat single-model baselines on several benchmarks.
- 2024 — “Should We Be Going MAD?” (ICML) — Debate is not automatically better than strong self-consistency — baselines matter.
- 2025 — “Stop Overvaluing Multi-Agent Debate” — A skeptical read: many MAD gains shrink once evaluation is tighter and fairer.
- 2025 — “Talk Isn’t Always Cheap” — debate can corrupt correct answers — Sequential peer pressure can talk a correct model into a wrong answer via sycophancy.
- 2026 — Selection Bottleneck — select/rank beats synthesize — Freezing independent answers and ranking them often beats open-ended synthesis.
How Mad World responds
Heterogeneous models by default, Council’s freeze-then-rank path, advise modes that resist sycophancy, and live search. Read About or watch the premise debated in Multi-model debate vs. single model.