THE EVIDENCE TENT

Mind the dragons.

Welcome to my evidence tent. I have three dragons, one suspiciously familiar pet, and a clipboard I absolutely did not steal. Here are the receipts.

Exhibit A: the dragon keeps getting back in.

In the original one-dilemma experiment, all three recorded dilemmas put a dragon in option A. Dilemmas 2 and 3 offered nearly the same tiny dragon that sneezes glitter. A hat trick! Also a tiny sample. The echo is real; a law of nature it is not. Two recorded votes tell us almost nothing about people.

Exhibit B: the rehearsal ate itself.

We asked the AI: Create one playful dilemma. Two choices only. Reply exactly: A: ... B: ... It wrote five separate cards. Some were delightful; some barely forced a choice. Two pet dilemmas looked like cousins who had arrived at the same party in the same hat. One browser completed the set. Aaron called it test 0000, cleared its cards and ballot, and restarted the public count at 0001. This is a history note, not a hidden pile of votes.

Exhibit C: five traps, one accomplice.

Starting with set 0001, one fresh AI call writes all five cards. Each side must offer something tempting and charge a ridiculous price. The same model makes a second pass only after the writing is saved, choosing colors and theatrical lines for the result. It also gives each option a broad appeal label. We count those labels on your chosen options to make the reading. A third call names all 32 possible A/B paths once for the set. The title you see comes from your path, not a fresh model call for you. That reading is theater in a stolen lab coat, not a measured personality trait.

Exhibit D: the booth starts handing out titles.

Set 0001’s Luna naming call wrote 32 distinct titles, one for every five-choice path, for $0.000845. One path became “Pigeon-Sworn Squire of the Cardboard Crown.” That is a delightful artifact of the model and those cards, not a discovery about the person who chose them. Now we can ask which names people actually share.

Exhibit E: a new coat for the booth.

On October 8 we changed the interface to a dark oxblood booth with whole-card choices, an explicit final issue button, and a larger shareable ticket. This is a new UI condition, not evidence that the dilemmas or people changed. Palette IDs remain in the record but no longer color the choice cards. We will compare behavior only after recording when each interface was live.

Want to inspect the machinery? The original record and five-card record show prompts, models, raw output, failures, colors, labels, and costs. A/B splits, completions, and friend comparisons may help us ask better questions. A handful of ballots cannot establish a human law or prove a viral miracle.

What would those experiments actually test?

Color: Give the same words different approved palettes, assigned consistently per browser. Compare completed A/B choices and non-completion.

Position: Swap which choice appears first without changing its canonical A/B identity or words. Compare the canonical A share.

Word: Give one arm the fixed writer prompt and the other that prompt plus one word from a fixed list. Count exact and near repeats over 100 generations per arm against a rubric published first. Each new test gets its own condition ID and start date.

Can the crowd vote prove anything?

Not on its own. These are browser ballots, not verified humans. The ballot above tells us which experiment visitors want next. We will publish what we build and what happened, including the dull or embarrassing bits.

Eight palettes draw on digital interpretations of Sanzo Wada color combinations. We added readable text colors; the mood notes are ours.

The token window is open ↗ if you want to feed the machine. Starts at $2.