Building Curio Experiment plan 2-minute read
Curio Will Compare Two AI Workflow Bundles on Fixed Briefs
Curio will compare two AI workflow bundles on fixed briefs and review rules. The test asks whether splitting idea work from building helps a solo builder.
The short version
- Curio will compare two workflow bundles under the same fixed briefs and review rules.
- The test matters because one weak direction can send a solo builder through the wrong build.
- Human creativity research separates making from judging, but it does not assign those jobs to AI models.
Curio will compare whole workflows
Curio plans to compare two AI workflow bundles on the same fixed briefs and review rules.
Leon Kelvin Li designs and builds Curio alone. There is no team handoff to catch a weak direction before it becomes a finished feature. A poor choice at the idea stage can send one person through a careful build of the wrong thing.
One bundle would use the proposed roles for idea work and building. The other would reverse those roles. The test asks whether the split helps the same solo builder choose a fitting idea and carry it into the finished work.
No result exists yet.
Human research separates making from judging
One human study used Israeli, South Korean, and Japanese samples. Participants rated ideas for traits such as originality and usefulness. In two studies, they rated ideas made by other people. In another, they rated their own.
The Israeli samples produced higher scores for varied and original ideas in the main group comparisons. They also rated ideas less strictly. Across the studies, rating strictness helped explain part of the difference in what the groups produced.
The authors said the two acts could not be fully pulled apart. A judgment can spark a new option, and a new option can expose a weak standard. The work studied people and culture, not AI.
Yvonne Görlich's self-report scale asks people about habits from finding a problem through testing, building, and sharing. It records how people describe their habits. It is not a map that every project must follow.
Curio's proposed routing is still only a plan
Leon proposes GPT-5.6 Sol for finding and judging early directions. After approval, a workspace option displayed as “Claude Fable 5 Max” would build the plan and run quality assurance, or QA. QA checks the work for missed rules, weak claims, and visible faults.
GPT-5.6 Sol is a documented OpenAI model. The Fable name is only a label in Leon's workspace; it does not reveal the model, provider, settings, or tools beneath it. That is the one naming limit the test cannot remove.
Past use did not hold the briefs, prompts, or tools still. It produced a routing idea, not evidence that either option is better suited to one role.
The experiment must score whole workflows
Curio can lock a set of briefs, source packets, tools, time limits, prompts, and review rules before the first run. One route would use the proposed roles. The other would reverse them.
Before the routes run, Curio can set the pass marks. A blinded reviewer would score idea fit without seeing the route. The reviewer would also count missed build rules. Time and cost would be recorded beside both results.
The swap compares two workflow bundles, including their order and prompts. It cannot isolate a model effect. Any result would apply only to the fixed briefs, tools, and rules used in the test.
For those briefs, the routing earns a place only if more ideas fit and the finished builds have fewer missed rules without taking more time or money. Otherwise, Curio should stop assigning work by the displayed label.