LoneStarCodeslonestar.codes
Field note · No. 2 · How we build

Design Thinking With Claude

The topologies and mechanics behind our multi-agent design teams — and a clean-room head-to-head against a single well-written prompt.

In our last post we made the argument for why multi-agent design works: a language model left alone gives you the modal answer, personas are a cheap way to sample from different regions of its distribution, and the value shows up when diversity is followed by adjudication. That was the theory. This is the practice manual — the actual mechanics and team topologies we run at Lone Star Observatories — followed by a live, clean-room experiment: two simple, universal design briefs, each given to a single directive prompt and to our iterative multi-team process. Every agent was a raw API call with no system prompt and no context of any kind, and the full unedited transcripts ship alongside this post.

Two results, up front, and one frame that holds them together: design is search over a solution space, and only one of these processes searches. A directive prompt buys you a point — one design, sampled from the center of the distribution: fluent, polished, and by construction the design everyone else asking the same question already has. Fluency is not design; the mode is where a search starts, not where it ends. The team buys you a map — the same space with its explored-and-rejected regions marked, its walls located, and its tradeoffs chosen on purpose. First result: the basin has real gravity — on the open brief, six independent samples landed in recognizably the same territory, down to the directive agent and the team converging on the same product name, and our own red-team agent said it to the team's face: "four people, one insight, dressed five ways. That's convergence, not creativity." Without structure, even conditioned samples fall back toward the average; a lone prompt never leaves it at all, because leaving requires the argument it doesn't contain. Second result: everything that makes the team's deliverable a design happened after the mode — five rejected regions with reasons marked on them, a wall nobody had noticed (weather), two ideas revealed to occupy incompatible coordinates so the synthesis had to choose, and an honest "this part of the map is unsurveyed." One process measures the prior. The other explores the space the prior sits in.

The frame: design thinking, mechanized

None of what follows is freestyle multi-agent enthusiasm — it is the classic design thinking loop, Empathise → Define → Ideate → Prototype → Test, run as a machine. Our squads execute it as Define → Ideate → Prototype → Test → Empathise → iterate-or-ship under a Facilitator who owns pacing and the iteration decision, with one deliberate twist: empathy comes around twice — as persona conditioning before anyone sketches, and as a formal radical-empathy phase after testing whose written output is explicitly allowed to contradict the original problem definition. The loop closes only when the Facilitator judges the definition has stopped moving.

F facilitator: pacing + the iterate/ship call DEFINE IDEATE PROTOTYPE TEST EMPATHISE iterate — the refined definition may contradict the first persona first-instincts argue what the problem IS blind parallel divergence → red-team convergence gate working artifacts · tournaments · subtractive charrettes identity swap: designers become . . the test users · snob boards radical empathy → refined problem definition (on disk)
Fig. 1 — The loop, mechanized. The classic framework with our machinery mapped onto each phase. Empathy appears twice: as conditioning before Define, and as the formal return leg after Test. The filesystem carries every phase's artifact to the next; the red iterate edge is the Facilitator's call, and it re-enters at Define — not at Ideate — because the thing testing most often breaks is the problem statement.

The strategies underneath everything

Five principles drive how we do design work with Claude. Everything else is implementation.

1

Design is search, not generation

A brief has a solution space, not an answer. One agent samples it once, near the center. We place multiple samples in different regions, then select. Divergence and convergence are separate mechanical steps — never one prompt.

2

Condition the samplers, don't heat them

Diversity comes from identity, not temperature. Every member gets a name, age, culture, and MBTI type. Identity reliably moves the reading of a brief; as the experiment shows, escaping the idea-basin takes the next principle too.

3

Make disagreement structural

"Consider other perspectives" is a sticky note. A roster where perspectives are assigned — one agent allowed only to attack — is a meeting where they show up and argue.

4

The filesystem is the memory

Every phase writes files the next phase reads. That's what makes iteration real: pass N pushes off pass N−1 instead of restarting from the same floor.

5

The human sits at the gate

No mid-session approval gates. The team runs autonomously; the Board sees the finished package — and the dissent that survived. Human attention is spent on verdicts, not supervision.

The five topologies

"Multi-agent" is not one architecture. We run five distinct shapes, chosen by the kind of question being asked. Filled sapphire nodes are facilitators/orchestrators; open nodes are persona agents; brick nodes are adversaries; the dashed node is the human.

F iterate until the facilitator is satisfied

1 · Facilitated round-table

The Design Thinking Squad: 4–8 personas around a Facilitator running Define → Ideate → Prototype → Test → Empathise → iterate. Every phase ends in a rendered team discussion, in voice, by name. Best for: "what should this even be?"

full-team debate at every gate intakestoriesanalysistestingeng.design looks like a waterfall; behaves like a gauntlet

2 · Pipeline with debate gates

The BA Team: Intake → Story Mapper → Analyst → Tester → Sober Engineer → Design Lead, with the whole team critiquing each deliverable before the next role begins. Best for: turning a settled direction into stories, acceptance criteria, and briefs.

human judges; zero cross-contamination

3 · Panel-and-judge

When the answer is an artifact, don't debate — compete. Five separately-conditioned teams shipped five complete Vision Pro concepts (Copernicus, Galileo, Halley, Hypatia, Tycho); selection deferred to the Board. Best for: directions where anchoring on the first concept costs the most.

claim "not proven" asserterred team cheapest topology per unit of error caught

4 · Adversarial pair

One agent asserts; a second — with standing to say "not proven" — attacks. Used small (a red-team pass inside a session, as below) and large (verification before anything ships). Best for: claims that matter.

BAUXdelivery rosters persist on disk — reload by name

5 · Hierarchy

A Programme Office of standing personas commissions design-thinking runs, turns output into business cases, and spawns BA/UX/delivery teams with briefs. The org chart is genuinely stateful, not a fresh cosplay per session. Best for: programmes, not projects.

Choosing between them

Open question? Round-table.
Settled direction? Pipeline.
High-stakes artifact? Panel-and-judge.
Claim that matters? Adversarial pair.
Many workstreams? Hierarchy.

And underneath every one of them: the filesystem, carrying the argument between phases.

Fig. 2 — The topology gallery. Five multi-agent shapes we run in production. Filled sapphire = facilitator/orchestrator · open = persona agent · brick = adversary · dashed = human. The dashed loops are debate/iteration paths.

The mechanisms that make the magic

Topologies are the org chart. The actual search moves — the things that reliably pull a session somewhere a lone prompt doesn't go — are three mechanisms we use everywhere, and each is a way of sampling a different region of the space on purpose.

1 · Weird-persona injection. A professional design team samples professional-designer space, so we deliberately seat people who have no business being in the room. When five of our teams took an astrophotography app's desktop pages through a design charrette, every team's core (facilitator, UX lead, domain analyst) was flanked by a professional astronomer, a professional photographer, a novice, and a space-obsessed teenager — and every pass was critiqued by a design-snob review board: a fine-art painter, an infographic designer, and an Apple-school web designer, none of whom owed the design anything. The credits are traceable in the reports: the painter is why every page's final form treats the image as the center of gravity ("everything else is a mat around it" was, verbatim, "every fine-art-painter critique, every page"); the infographic designer contributed the honesty spine — every number carries its reference frame; every threshold is published and editable. The novice and the teenager are boundary probes: they ask for the thing professionals forgot to want and refuse the thing professionals forgot is confusing. A weird persona is a cheap coordinate transform — the same model, teleported to a far region of the space, reporting back what the design looks like from there.

2 · Identity swapping. The cheapest usability lab in existence: at the Test phase, the design team becomes the test users. The instruction, from a real session's facilitator, is one line: "Step out of yourselves. Be the tester. Be honest, including when it hurts." Agents who spent three phases advocating for the design are re-conditioned into the target personas from the test plan and made to meet their own prototypes cold. It works for the same distributional reason everything here works: an identity is a position in the space, and swapping identities moves the observer instead of the artifact — the designer-persona knows what the button is for; the user-persona it becomes genuinely doesn't, and says so, in character, with feelings. It's also what keeps persona conditioning from hardening into persona lock-in: nobody keeps one seat for a whole session. The Empathise phase then formalizes what the swap surfaced into the refined problem definition.

3 · The subtractive charrette (remove-a-component). Generative models accrete — ask for a refinement and you will overwhelmingly get additions. Deletion is a search operator the model almost never applies to its own output, so we force it structurally: in the charrette above, five teams passed each page around a Latin square — every team touched every page — and each pass ended with a hard rule: delete one fresh major component and hand the blank to the next team, who must decide what best deserves the reclaimed space. Thirty agents, thirty files, zero re-used deletions. The same class of component kept getting cut and never missed — activity logs, KPI walls, storage bars, standing settings forms — and the through-line the deletions revealed became design law: most "panels" are an intermittent need masquerading as a standing one; only things that change while you watch earn permanent pixels. The final report's summary is the mechanism's whole defense: "the page got quieter and more honest at every step, and never lost a capability — each deletion was a relocation, not an amputation."

ONE PAGE · FIVE PASSES (Latin square: every team touches every page) PARALLAX depth & context spec → add → refine ✂ delete 1 component blank AIRMASS ruthless subtraction fill the blank → refine ✂ delete (a fresh one) AVERTED VISION progressive disclosure fill → refine ✂ delete CULMINATION peak-moment focus fill → refine ✂ delete ALBEDO legible at 2 a.m. fill → refine → v5 owner writes REPORT EVERY TEAM'S ROSTER core: facilitator · UX lead · domain analyst + pro astronomer · pro photographer · novice · space-obsessed teenager ← the weird-persona injection THE SNOB REVIEW BOARD · critiques every pass fine-art painter — "the image is the centre of gravity" infographic designer — "every number carries its reference frame" Apple-school web designer — restraint, hierarchy, one lit action
Fig. 3 — The subtractive charrette, as actually run. Five teams × five pages, 30 agents, 30 files, zero re-used deletions. Each pass: full component spec → brainstorm additions → snob review + revisions → delete one fresh major component → hand the blank on. The deletion rule is the anti-accretion engine; the weird roster and the hostile board are the coordinate transforms. Five versions per page, then the owning team writes the recommendation.

Map these back to the framework and the shape of the whole system becomes visible: weird personas widen Ideate, identity swapping powers Test and feeds the Empathise return leg, and the subtractive charrette is what iteration looks like when you refuse to let it mean accretion.

The outputs we demand

A topology without deliverable discipline produces vibes. Every phase has a named artifact with a schema, and the artifacts are designed to be arguable with — they preserve disagreement and provenance rather than laundering everything into consensus prose. This is what one real session directory looks like:

designthinkers/<team>/<session>/ ├── 01-problem-definition/ │ ├── mission-statement.md "[user] needs [need] because [insight]" │ └── how-might-we.md 4–6 HMW prompts + assumptions list ├── 02-ideation/ │ ├── raw-ideas.md 5–8 ideas each, every idea attributed │ └── shortlist.md scored on 4 criteria, with rationale ├── 03-prototypes/ │ ├── prototype-<name>.html working single-file, clickable │ └── prototype-<name>.md concept/storyboard for non-UI ideas ├── 04-testing/ │ ├── test-plan.md persona + what failure looks like │ └── results.md simulated sessions, cold reactions ├── 05-empathy/ │ └── refined-definition.md allowed to contradict 01 — the point └── final-presentation/ ├── presentation.html stakeholder-grade deck └── summary.md + provenance per element
Fig. 4 — The artifact contract. The session directory from our Design Thinking Squad skill. Because every idea stays attributed all the way through, we can audit where a final design came from — which is how we can tell you, below, which elements came from convergent evidence, which from a single voice, and which from the critic.

The experiment

Claims like these deserve a test you can inspect, so we ran one under clean-room conditions: two simple, universal briefs; every agent a raw API call to the same model with no system prompt — the prompt is the entire context; persona prompts specify identity, role, and schema only; downstream agents receive prior outputs verbatim, not summarized; one pass, no cherry-picking. Full transcripts with per-call token counts: every agent's unedited output.

CONDITION A — DIRECTIVE one agent,one prompt concept doc 29s · 1.7k tokens CONDITION B — DIVERGE → ATTACK → SYNTHESIZE Ingrid · ISTJ Dayo · ENTP Mercedes · INFP Ken · ENTJ parallel, blind to each other Priya · red team receives all four, verbatim F facilitator synthesis concept doc + provenance · 90s · 18k tokens
Fig. 5 — The experiment design (brief 1 shown, with its measured costs). Both conditions receive the identical brief and deliverable schema, on the same model, in empty contexts — a clean-room slice of the loop's Define → Ideate → converge leg. Brief 2 uses the same shape with three drafters and no separate red-team stage.

Brief 1 (open-ended): the first night

"Design an experience around a person's first night out with their first telescope."

The directive agent produced "First Light" — kit plus app, one promise: "you will see something breathtaking tonight, within 30 minutes of stepping outside." It is fluent and confident: daylight assembly with a finderscope-alignment checkpoint, one "hero object" instead of a menu, a deliberate 60 seconds of interface silence at the eyepiece, a second-session metric, even a quotable thesis ("the product's most important design decision is knowing when to disappear"). Hold that inventory in mind — every one of those moves is about to reappear, unprompted, in the blind samples below. Nothing here was found. It was recalled.

The four persona ideators, running blind, read the same brief four different ways:

ISIngrid, 58ISTJ · ex-engineer
DADayo, 29ENTP · reframer
MVMercedes, 41INFP · UX researcher
KNKen, 35ENTJ · growth lead
PRPriya, 47INTJ · red team only

"The first night is not an experience problem — it's a failure problem… Saturn's rings sell themselves. Frustration is what needs engineering."

Ingrid · first instinct

"Don't design the first night with a telescope. Design the first night with the sky, where the telescope earns its entrance."

Dayo · first instinct

"Underneath all of it is this fragile, embarrassing hope: I wanted to feel wonder tonight, and instead I feel stupid… The person's relationship with their own hope is the product."

Mercedes · first instinct

"Let's be honest about what this brief actually is: a churn problem wearing a romance costume."

Ken · first instinct

Notice what happened: the readings genuinely diverged — commissioning, expectation, shame, retention — but the ideas underneath kept converging: everyone wants a daylight rehearsal, a single guaranteed target, a designed second night. The directive agent, with no persona at all, had found most of the same moves. Five samples, one basin. The red team is the agent that noticed — and Fig. 6 shows what she did about it.

Ingrid
→ CORE
Daylight rehearsal — align the finder before dark
→ kept
Date-keyed target card with a sketch of what you'll actually see
→ kept
Comfort checklist · laminated fix sheets on the mount
→ kept
"Design the second night, not the first"
Dayo
→ kept
Court-the-sky-first sequencing; "blob appreciation" reframe
✕ killed
The Locked Cap
"clear skies are perishable; burning them on principle is design vanity"
✕ killed
Twin First Light (paired strangers)
four dependencies, zero in our control
✕ killed
Disappointment Postcard
Mercedes invented it too — "convergence, not creativity"
Mercedes
→ CORE
Protect hope, not features; whisper companion (audio, no screen)
→ kept
One Guaranteed Win; re-entry ritual ("tell someone what you felt")
✕ killed
Postcard to yourself
duplicate; deferred payoff doesn't fix immediate churn
Ken
→ kept
One-target-no-menu guarantee; funnel framing
→ kept
48-hour hook; second-session KPI
✕ killed
Shareable proof-of-win badge
"Ken designed the funnel and forgot the dark"
Priya · red team
+ found
"Weather. Nobody said the word." — the cloudy first night
+ found
The Moon isn't always up — the guarantee evaporates
+ found
Scope of control unstated (app? packaging? hardware?)
+ found
Groupthink: "four people, one insight… all anecdote, zero research"
Synthesis → "First Light: a two-night, screen-dark protocol"
  • Scope stated plainly: printed in-box kit + optional audio companion — what a telescope maker actually controls
  • The Ken/Mercedes contradiction resolved: the phone exists before sunset and after pack-up only — audio and red-light print at the eyepiece
  • Clouds get a designed path: a cloudy first night is a rehearsal night — "waiting is the hobby's first skill"
  • The single target becomes a date-keyed ranked ladder (Moon → planet → double star → cluster) — the promise never rests on one object
  • Daylight rehearsal as persuasion, not a padlock: "skip this and tonight will fail"
  • Ten minutes of naked-eye orientation first — Dayo's instinct, minus his hostage mechanism
  • 48-hour hook and second-session-within-7-days as the north-star metric
  • And an admission: the expectation-gap thesis is anecdote — diary studies with 20 real first-nighters before scaling
Fig. 6 — Idea provenance, divergence to synthesis. Struck-through chips were killed by the red-team pass, with her reason attached; her column contains only findings, because her prompt forbade her from ideating. Dot colors on the final concept credit each element's origin. All text is quoted or condensed from the raw transcripts.
UNBOXING eyepiece photos DAYLIGHT REHEARSAL "skip this and tonight fails" THE PLAN phone · daylight CLEAR? yes ONE-TARGET NIGHT screen-dark · whisper-led ranked target ladder 48h hook no CLOUDY FIRST NIGHT PATH "waiting is the hobby's first skill" phone allowed here… …never here
Fig. 7 — The synthesized journey. The two structural differences from the directive design are both map features, not corrections: the weather branch marks a wall of the space the red team noticed nobody had surveyed, and the screen rule ("the funnel ends at the back door and resumes at breakfast") marks a fork the design now straddles knowingly instead of accidentally. The directive design — whose promise is "you will see something breathtaking tonight" — stands on the same terrain without the map: no cloudy branch at all.

Brief 2 (basic): a page about a star

"Design a web page that shows information about a single star — take Betelgeuse as the example."

The three drafters split on what the page is for. Ruth (ISFJ): the real visitor arrives from a search — "is Betelgeuse going to explode?" — anxious about looking foolish; answer their question first, in plain words. "Clarity is kindness." Beto (ENFP): this is a deathwatch — the page should leave you "small, lucky, and a little haunted." Anja (INTJ): comparators and honest uncertainty are the page — distance shown as a ±90-light-year band, mass as a range ("show the range, not a fake midpoint"), a live light curve because a static magnitude misleads, and a "How we know" section: "provenance is what separates reference from decoration." The facilitator named the real disagreement — what the first ten seconds promise: reassurance, awe, or rigor — fused all three into one hero, and killed Beto's "could explode before you finish reading this page" hook with the sharpest sentence of the run: manipulating the anxious reader we exist to reassure is "a trust debt no engagement metric repays."

Condition A · directive, one agent

"The design becomes the data" — a variability-pulsing portrait.

example.com/stars/betelgeuse
Betelgeuse

A dying giant, 548 light-years away.

~548 lyM1-Ia~764 R☉Orion
Scale — slider: Betelgeuse swallows the inner solar system
The Great Dimming — scrollytelling; the page itself dims
Life & Death — supergiant stage, impending supernova
Sky Finder — find it in Orion tonight

Condition B · team synthesis

The question answered in the hero; uncertainty as first-class content.

example.com/stars/betelgeuse
This star will explode — sometime in the next 100,000 years.

Nothing dangerous. Everything wonderful. 🔊 BEET-l-jooz

Distance: 548 ± 90 light-years

The light in your eye left it around the time the Ming dynasty ruled. The band is the honest answer — parallax is hard for a star this bloated.

↻ flip: how we measured this
the Great Dimming · 2019–20
how big is dying? · what you'd actually see · find it tonight · vital signs (ranges, museum-placard style) · how we know
Fig. 8 — The same brief, both answers, rendered. Both heroes pulse on the star's variability — the directive agent and the team's narrative designer invented that move independently, in empty contexts. Two clean samples, same flourish: that is the center of the distribution announcing itself, and we report it as a phenomenon, not a prize. What only the team produced is the trust layer — the famous question answered plainly, ranges instead of point values, the light curve, "How we know" — and a documented rejection of a manipulative hook.

The return leg, run clean-room: empathy, randomness, subtraction, and the sun

The two briefs above exercised only the Define → Ideate slice, and we said so. So we ran the rest of the loop under the same clean-room conditions. The two first-night concepts from brief 1, verbatim, met four identity-swapped user personas, each living a different real night: Marcus (34, two kids with twenty minutes of patience, clear with a half Moon), Eleanor (61, app-averse, completely clouded out), Priya (24, anxious perfectionist, moonless sky), Tyler (27, phone-native, under a parking-lot light dome). The directive concept was frozen — one pass is its definition — and tested once. The team concept entered the house loop: each iteration ran the empathy round (four blind persona nights), injected one outside voice from a rotating weird deck, and ended in a facilitator revision with a mandatory deletion and a refined definition allowed to contradict the original. And the iteration count was tied to the sun: a fixed 420-second night clock. Nobody chose the number of iterations. The clock did.

SUNSET t+0 · v1 meets four strangers SUNRISE t+420s v5 ships untested v2 t+79s · lighthouse keeper ✂ naked-eye orientation re-test v3 t+191s · kindergarten teacher (the fire drill · name blank) ✂ the Whisper Companion — "silence becomes the default state" re-test v4 t+296s · stage magician (the Prediction Envelope) ✂ the Star-Hop panel — which v3 itself created re-test v5 t+415s · air-traffic controller (the Go-Around Card) ✂ Cloudy card → generalized into scripted exits for every failure mode every version: 4 identity-swapped nights + 1 outside voice + 1 mandatory deletion · the directive concept sits frozen at t+0 by definition
Fig. 9 — The night clock. Four iterations happened because 420 seconds of darkness allowed four — the count was the clock's choice, not anyone's. Sunrise arrived five seconds after v5 was produced, so v5 ships untested. That is the stopping rule doing its job: "are we done?" is a question the mode answers instantly and iteration, left to itself, answers never; the sun replaces it with one that always has an answer — what does the darkness we have left buy?

Eleanor's clouded-out night is the whole argument in one persona. Against the directive design she wrestled a QR code at three distances and received: "Cloud cover 100%. No targets visible tonight." Full stop. Her verdict — yes to a second night, "but because I've waited since I was ten and the clouds will break, not because this design earned it; it planned one perfect night and forgot that weather, knees, and grandmothers exist." Against the team design she found the Cloudy Night path on paper, ran the daylight rehearsal anyway, and had her first light in the afternoon — a pigeon on a chimney pot: "The finder agreed with the eyepiece because my hands made them agree… The clouds took the sky, not that."

An honest caveat before the score: all twenty simulated nights ended in SECOND NIGHT: yes — synthetic users are agreeable, and we won't pretend otherwise. The signal lives in the clause after the yes. The directive design's yeses arrive despite it; the team design's arrive because of named mechanisms: "only because the fix sheet made failure survivable""the design let me fail privately, fix it myself, and still win""when it broke, the paper let me fix it in front of my kids instead of failing in front of them."

And the loop ate its own darlings, on evidence. v3's deletion was the Whisper Companion — the audio guide the original synthesis was proudest of — with the ruling that "its best feature — the silence — becomes the default state of the whole design." v4 then deleted the Star-Hop panel v3 itself had just created: the loop correcting its own previous iteration, something a frozen artifact cannot do by definition. The weird deck earned its seats line by line — the lighthouse keeper's assemble it blind once, the kindergarten teacher's fire drill ("introduce heroes before the crisis"), the magician's sealed envelope ("people return to collect promises"), the controller's pre-committed aborts ("a go-around never feels like quitting, because we rehearsed it as success"). The shipped v5 opens: "Paper-only, any-sky, two-night protocol… Every session ends on a printed card, never on a shrug." Cost of the whole return leg: 28 calls, 105,085 tokens, one 420-second night — every addition tagged with the persona or voice that earned it.

Scoring it honestly

Total tokens, measured (relative to directive)
Brief 1
1× (1.7k)
10.6×
Brief 2
1× (0.8k)
6.5×
Wall-clock time, measured (seconds)
Brief 1
29s
90s
Brief 2
16s
33s

Directive (1 agent)Team (4–6 agents)

Fig. 10 — What the map costs, from the API's own usage numbers. Parallel divergence keeps wall-clock sub-linear in agent count — six agents cost ~3× the time, not 6×. Tokens scale super-linearly because downstream agents re-read everything verbatim. That re-reading is the surveying; it's what you're buying.
Directive (1 agent)Team (4–6 agents)
Brief 1: core ingredientsDaylight setup · one target · silent moment · 2nd-night metricSame ingredients, reached independently
Brief 1: weatherAssumes a clear night — the promise breaks on cloudsNamed as blind spot; cloudy-night path designed
Brief 1: screens in the darkApp at the eyepiece (red mode + photo prompt)Contradiction named, resolved: phone banned after sunset
Brief 1: assumptionsShipped unexamined"All anecdote, zero research" — validation step added
Brief 2: signature moveVariability-pulsing pageSame pulse, invented independently — the modal move
Brief 2: trust layerNoneAnswered question · uncertainty bands · "How we know" · hook killed for cause
Killed ideas w/ reusable reasons05 (+4 mandatory deletions in the return leg)
Eleanor's clouded-out night"It planned one perfect night" — dead on arrivalCloudy path + daylight first light — "the clouds took the sky, not that"
Second-night yeses (all 20 said yes)4 — despite the design16 — because of named mechanisms
Iterations1, by definition4 — the count chosen by the sun
Self-correctionNone possiblev4 deleted what v3 created
ProvenanceNoneEvery element attributed, incl. the weird voices
Fig. 11 — The verdict table. One frame throughout: the mode is a magnet — strong enough that even persona conditioning barely left its neighborhood, which the team's own critic diagnosed live. A process that returns the unexplored average has not designed anything; it has measured the prior. The single prompt is the null hypothesis, not a rival. What only the search produces is the map: rejected regions with reasons, located walls, forks chosen knowingly, unsurveyed terrain flagged. Independent divergence draws the basin's extent; assigned adversariality walks its boundaries; synthesis with provenance marks the route.

A directive prompt buys you a point. The team buys you the map.

What we are not claiming

Same disclaimers as last time, because they still hold. This is orchestration, not new capability — every call in every condition hit the same model. Two briefs plus one return-leg run is an illustration, not a study, and the team condition consumed an order of magnitude more inference. The simulated users all voted yes — synthetic testers are agreeable, which is why we scored the reasons, not the verdicts, and why real diary studies remain in the shipped concept's own plan. Prototyping is the one phase of the loop not demonstrated here. And the directive baseline is not a strawman — which is precisely the hazard worth naming: its outputs are fluent, plausible, and heavily overlapped with the team's instincts, and nothing in the artifact discloses that no search happened. The unexplored average arrives wearing the confidence of a finished design. That is why "it looks done" is the least trustworthy property a design can have.

One more disclosure, in the spirit of this series: our first attempt at this experiment ran as subagents inside our own workbench, and a diagnostic probe showed those agents were inheriting our organization's full project context — they were not the blank agents the comparison required. We threw that run away and rebuilt everything as raw, context-free API calls. The failure mode is worth naming because it is the multi-agent version of the modal trap: your agents are only as independent as the context you didn't notice they share.

We keep it dark out there by resisting the easy, probable, beige answer. Sometimes that takes six agents and an argument. Sometimes it takes one agent and the discipline to leave it alone. Knowing which brief is which is the actual skill.

Full unedited transcripts, per-call token counts, and run metadata: transcripts.md. Every call in both conditions ran on claude-fable-5, 2026-08-11, single pass, no system prompt. Diagram palette validated for color-vision-deficiency separation and contrast in both light and dark themes.

← All posts