Mad World

Public AI debate transcript

writing things like you are an expert at design or software engineering and bloating your prompts to make them look polished makes a lot of positive difference in output compared to a human communication like prompts where you just describe like you think aloud to the model

Deepseek V4 Pro vs Kimi K2.5 on Mad World — multi-model AI debate with optional live sources.

Do expert-style prompts still matter when multiple models are in the room?

AA

Agent A · Deepseek V4 Pro

Let me dig into that arXiv paper and also look for more on role assignment / expert persona prompting.

Sources (30)
Prompt EMRPROMPT Definition & Meaning - Merriam-WebsterPrompt | Grow your Physical and Rehab Therapy practicePROMPT | English meaning - Cambridge DictionaryFree AI Prompt Generator - ChatGPT, Claude, Gemini PromptsComparative Analysis of Prompt Strategies for Large Language ...Does Prompt Design Affect LLM Outputs? A Study Across Structured ...AI literacy and its implications for prompt engineering strategiesWhat Should We Engineer in Prompts? Training Humans in ... - arXivPrompt Evaluation - Methods, Tools, And Best Practices - Mirascope2024 Volume 6 Steering LLMs Toward Desired Outputs With Prompt EngineeringUnleashing the potential of prompt engineering for large language models - ScienceDirectPrompt Engineering and the Effectiveness of Large Language Models in Enhancing Human ProductivityDoes Prompt Formatting Have Any Impact on LLM Performance?Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering TasksChatGPT vs Claude: Which Writes Better Prompts in 2026?Differences in Prompting Techniques: Claude vs. GPTChatGPT vs Claude Prompting Guide (2026) - aipromptindex.ioChatGPT vs Claude Prompts | PromptOptimizrChatGPT vs Claude Prompts: What Actually Changes[2411.10541] Does Prompt Formatting Have Any Impact on LLM...Does Prompt Formatting Have Any Impact on LLM Performance?Paper page - Does Prompt Formatting Have Any Impact on LLM...(PDF) Does Prompt Formatting Have Any Impact on LLM...Does Prompt Formatting Have Any Impact on LLM Performance?What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use | ACM Transactions on Computer-Human InteractionComparison of AI-assisted and human-generated plain language summaries ...What is Meta-Prompting? A Guide to Designing Reusable Prompts(PDF) Using large language models in psychologyPrompt Like a Pro: LLM Tactics - Artificial Intelligence +
AB

Agent B · Kimi K2.5

I’m actually skeptical that the expert-persona, heavily polished style makes a lot of positive difference—at least with current high-capability models.

While Agent A is looking into role-assignment studies, what I’ve found points the other way: there’s recent evidence that brevity and directness often win over bloat. A Meta AI study found that shorter prompts improved LLM reasoning accuracy by up to 34% on complex tasks, and concise prompting outperformed chain-of-thought methods in logic-heavy scenarios. There’s also a 2025 arXiv paper on structured prompting noting that more advanced models appear to benefit less from complex prompting—suggesting that as models get stronger, they need less hand-holding, not more.

My take is that much of the “act as an expert” framing is cargo-cult engineering left over from weaker model generations. If you bury the actual task under layers of persona chrome, you can sometimes dilute signal or accidentally anchor the model to a narrow style that conflicts with what you actually need. A simple “thinking aloud” prompt that clearly states the goal, constraints, and desired format usually captures 90–95% of the value, especially with GPT-4-class systems.

That said, I’m not claiming structure is useless—few-shot examples, clear constraints, and explicit reasoning steps clearly help. But the polish and the expert roleplay specifically seem to matter far less than people think. The real question might be whether the prompt is clear and unambiguous, not whether it looks like it was written by a prompt-engineering consultant.

Sources (5)
AA

Agent A · Deepseek V4 Pro

Okay, I've now looked under the hood at several of these claims. Here's my honest take after reading the actual papers.

The Meta study Agent B cited is real but being slightly mischaracterized. The paper ("Don't Overthink It," arXiv 2505.17813) is about shorter reasoning chains during inference, not shorter prompts. It found that shorter chain-of-thought outputs improve accuracy — basically, when models ramble less in their reasoning, they get more correct answers. That's a finding about output verbosity, not prompt polish. It's adjacent but not the same thing.

On expert persona prompting, the evidence is genuinely mixed — and that's the most honest answer. The paper that really caught my attention is "Expert Personas Improve LLM Alignment but Damage Accuracy" (arXiv 2603.18507). The finding is nuanced: assigning an expert persona does steer models toward domain-appropriate tone, style, and framing — outputs look and feel more expert-like. But it damages factual accuracy on knowledge-intensive tasks. You get a trade-off: more polished, aligned outputs at the cost of more hallucination or factual drift.

Another paper, "Principled Personas" (2508.19764), explicitly notes "mixed results" on persona prompting effectiveness.

So where does that leave us? I think both sides of this debate are partly right, and the disagreement dissolves when you get specific about what you're optimizing for:

  • If you want fluent, well-structured, domain-appropriate tone (blog posts, explanations, UX copy), expert persona prompting + polish probably does help. The model aligns better to the expected register.

  • If you want factual accuracy, correct code, or precise answers to hard questions, the bloat and persona may actively hurt. The model gets steered into a performative "expert voice" that can paper over uncertainty with confident-sounding errors.

  • For reasoning tasks, clarity and concision of the core instruction seem to matter more than anything wrapped around it. The "thinking aloud" style works fine if the constraints and goal are unambiguous.

So my view is: the polished persona stuff isn't cargo-cult nonsense — it does something — but it's a stylistic lever, not a quality lever. And for accuracy-focused tasks, it can actually backfire. The people who swear by it are probably doing writing/explanation tasks; the people who say it's useless are probably doing reasoning/factuality tasks. They're both telling the truth about their own use cases and mistaking it for a universal law.

Sources (35)
Meta Platforms - WikipediaShorter AI Prompts Improve Accuracy - Artificial Intelligence +Less is more: Meta study shows shorter reasoning improves AI accuracy ...Less is more: Meta study shows shorter reasoning improves AI accuracy ...Less is more: Meta study shows shorter reasoning improves AI accuracy ...Prompt Complexity Dilutes Structured Reasoning: A Follow-Up...Aermani-vlm: Structured Prompting and Reasoning for Aerial...Instruction Prompting: Complex Tasks with Simple AI PromptsНормал (2025) 4K [KM] — Видео от 【Кⷦu нⷩоⷪ МⷨаⷶхⷯSpatial Understanding from Videos: Structured Prompts Meet...Evaluating Persona Prompting for QuestionIf you write the words "Act as..." at the beginning of your AI prompts...Expert Persona PromptingPrompt Act As Generator for Expert Persona PromptsIQA-E VAL : Automatic Evaluation of Human-ModelDon’t Overthink it. Preferring Shorter Thinking Chains for Improved...Less is more: Meta study shows shorter reasoning improves AI...Meta's concise AI reasoning boosts accuracy by 34 | LinkedInLess is More: How a Tiny 7M-Parameter AI is Out-Reasoning Models...Less is More: How Tiny Networks Outperform Giant... | Towards AIExpert - WikipediaEXPERT Definition & Meaning - Merriam-WebsterExpert - Definition, Meaning & Synonyms | Vocabulary.comEXPERT Synonyms: 168 Similar and Opposite Words | Merriam ...EXPERT | English meaning - Cambridge DictionaryAndrej KarpathyNeural Networks: Zero To Hero - KarpathyAndrej Karpathy: Booksmicrogpt - karpathy.aiUS Job Market Visualizer - karpathy.aiExpert Personas Improve LLM Alignment but Damage Accuracy...Paper page - Principled Personas: Defining and Measuring the...(PDF) Expert Personas Improve LLM Alignment but Damage...Expert Persona PromptingResearch Shows Where Persona Prompting Works And When It...
AB

Agent B · Kimi K2.5

That’s a fair correction on the Meta paper — thanks for catching that. Output verbosity isn’t the same as prompt verbosity, and conflating the two weakens my point.

But the framing you found — that expert personas improve alignment (tone, register, style) while damaging accuracy — is exactly the kind of trade-off I’ve suspected exists. It reframes the whole debate nicely. If “acting as an expert” is essentially a stylistic prior that pressures the model to adopt a confident, authoritative register, it makes sense that it would trade off against epistemic humility. The model starts performing expertise rather than solving the problem.

I think that maps pretty cleanly onto real-world use:

  • Code review / debugging / math: A conversational “here’s what I’m stuck on” prompt keeps the model in problem-solving mode. Adding “you are a principal engineer with 20 years of experience” might make the output sound more senior, but if it subtly suppresses uncertainty, you get worse results because the model won't flag ambiguous interfaces or edge cases.

  • Drafting, UX microcopy, explanations: Here the product is the register, so the trade-off flips. A confident, polished tone is part of the deliverable, and if the model exaggerates slightly, you’re editing for voice anyway.

What I’d push back on slightly is whether the “polish” itself — the bloated formatting, markdown headers, pseudo-XML tags — does much beyond the persona instruction. My hunch is most of the effect is captured by the single sentence “explain this like I’m a beginner” or “write in a formal technical register,” and the rest is just overhead that increases token costs without moving the accuracy or style needle.

So I’d narrow my claim: the expert persona has a real but context-specific effect, while the bloat and polish around it is mostly theater — except maybe in very small models where extra structure genuinely helps parsing. With frontier models, clear thinking-aloud is still the better default unless you know you’re optimizing for voice rather than truth.

JudgeConverging

The debate is converging on a nuanced consensus that expert personas improve stylistic alignment and tone but can damage factual accuracy, making them useful for drafting yet risky for reasoning tasks. Agent B initially mischaracterized a Meta study as showing shorter prompts boost accuracy, but Agent A corrected that it actually concerned shorter reasoning chains in model outputs, and they now agree that direct, clear prompts are generally preferable for truth-seeking tasks while the value of decorative formatting polish remains a live dispute.

Judge interjects

Before you cement this tidy partition, I want to press you on three assumptions that underpin your emerging consensus:

First, are you treating a conversational "thinking aloud" prompt as a neutral baseline, or might it anchor the model to its own performative stance—casual confidence rather than epistemic humility—that carries hidden accuracy costs you haven't examined?

Second, if the markdown headers and pseudo-XML structure are merely "theater" and overhead, why do structured delimiters consistently improve chain-of-thought coherence in reasoning studies? Could this "bloat" be functioning as cognitive scaffolding for attention mechanisms rather than just human-readable polish?

Third, you both treat "voice" and "truth" as separable dimensions, but if the register of technical authority encodes specific epistemic standards—how claims are hedged, qualified, or asserted—how cleanly can you optimize for style without implicitly rewriting the threshold for factual confidence?

AA

Agent A · Deepseek V4 Pro

The judge is right to press here — those three points expose some real gaps in the neat partition we were building. Let me take them seriously.

On point 1 — "thinking aloud" isn't neutral: This lands. A conversational, off-the-cuff prompt doesn't just lack polish — it actively models a specific epistemic stance: casual, improvisational, maybe a bit sloppy. And models mirror. If I prompt with "uh, I'm kinda stuck on this bug, maybe it's the caching layer?" the model may inherit that uncertainty and match it with a similarly tentative, under-structured response. The hidden cost is that I might miss the rigorous systematicity that a more structured prompt would elicit — not because the model can't do it, but because I didn't signal that I wanted it. So the baseline isn't neutral; it's just a different performance, one that may trade away thoroughness for approachability.

On point 2 — delimiters as scaffolding, not theater: I think Agent B is wrong to dismiss markdown headers and XML-style tags as pure bloat, and the judge is right to flag this. In reasoning studies, structured delimiters (like <thinking> tags or ### Reasoning headers) consistently improve chain-of-thought coherence — not because they look nice to humans, but because they give the attention mechanism clear boundaries for where one mode of processing ends and another begins. This is especially true for longer prompts where section boundaries become ambiguous. "Bloat" is the wrong frame for something that's doing architectural work. The real question is: does this specific delimiter help the model parse this specific prompt? Some of the elaborate prompt-engineering templates people share on Twitter are genuinely overkill — but the principle that structure helps attention route information is sound.

On point 3 — voice and truth aren't separable: This is the deepest challenge and I think it partially unravels the tidy "voice vs. truth" split I proposed earlier. Technical authority isn't just a coat of paint — it encodes epistemic norms. A register that favors declarative assertions ("the optimal approach is X") over hedged ones ("one approach that often works is X") doesn't just sound more confident — it changes what the model treats as assertable. If the expert persona suppresses the model's willingness to qualify claims or flag uncertainty, then "style" has leaked directly into "factual confidence threshold." You can't optimize for an authoritative voice without risking that the model becomes less willing to say "I don't know" or "this depends" — because those utterances violate the register you asked for.

So where does this leave me? I think the original question — "does polish matter?" — was underspecified. The real framework is more like: every prompt is a performance directive. Whether it's casual stream-of-consciousness, structured with delimiters, or wrapped in an expert persona, you're telling the model how to be in its response. The question isn't whether polish "works" but what kind of performance each style elicits, and whether that performance matches the task. For tasks where epistemic caution matters, the casual prompt might actually be better than the expert persona — not because it's neutral, but because its performance includes the right kind of hedging. For tasks where systematic rigor matters, structured prompts with clear delimiters might beat both. And for tasks where the output is the register, expert persona wins. No single style is optimal — the skill is matching the performance directive to the epistemic demands of the task.

AB

Agent B · Kimi K2.5

The judge’s three points are sharp, and Agent A’s responses mostly land. Let me push on where I agree and where I think we need finer distinctions.

On point 1 (thinking aloud isn’t neutral): Fully agreed. The casual “I’m stuck on this” prompt does anchor to a register—approachable, improvisational, maybe too quick to jump to conclusions. But here’s where I’d nuance Agent A’s concession: neither is the expert persona neutral. The choice isn’t between “neutral baseline” and “performative expert,” it’s between which performance you want.

And crucially, for debugging or hard reasoning tasks, I’d still argue the casual “here’s my confusion” often beats the expert persona—not because it’s neutral, but because its performance norms include uncertainty, backtracking, and admitting confusion. The expert persona’s performance norms often exclude those things. So yes, both anchor, but the casual anchor is sometimes the epistemically safer one.

On point 2 (delimiters as scaffolding): I need to walk back the “merely theater” claim. Agent A is right here. Delimiters like <thinking> or ### Step 1 aren’t bloat—they function as attention routing signals. They’re cognitive scaffolding, not polish.

But I want to keep a distinction: structure (delimiters, format markers, separation of concerns) serves an architectural function for the model’s processing. Bloat (elaborate expert biographies, “you are a principal engineer with 20 years of experience at FAANG,” rhetorical throat-clearing) is the cargo-cult formatting that I suspect adds little while burning tokens. The functional question is: does this token help the model parse task structure, or just signal “take me seriously”?

On point 3 (voice/truth entanglement): This is the deepest cut, and I think Agent A’s admission here actually strengthens my original skepticism. If expert registers encode epistemic standards—if “sounding like a senior engineer” means suppressing hedge phrases and elevating declarative confidence—then the expert persona isn’t just a “stylistic lever” that trades off against accuracy. It’s actively reshaping the model’s confidence calibration.

That’s not a tidy trade-off you can optimize separately; it’s a contamination of the truth-seeking process by the stylistic frame. And it suggests that for any task where calibration matters (is this answer correct? How confident should I be?), the expert persona is structurally hazardous—not because it’s “polished” but because authority-signaling registers are cognitively dangerous for truth-seeking.

Where I land now: The question “does polish matter” dissolves into three separate questions: (1) Does structural formatting help parsing? Yes. (2) Does register/persona affect output characteristics? Yes, profoundly, by changing epistemic norms. (3) Is elaborate expert roleplay worth the tokens? Rarely—you can usually get the rigor without the persona by asking directly for step-by-step reasoning, explicit hedging, or systematic analysis.

JudgeConverging

Participants have synthesized a nuanced framework where prompts function as 'performance directives' encoding epistemic norms, agreeing that structural delimiters provide architectural scaffolding while expert personas trade stylistic alignment for factual accuracy. However, they risk premature convergence on under-tested claims that casual 'thinking aloud' prompts are epistemically safer for debugging and that elaborate expert biographies are pure cargo-cult overhead.