Hi Connor,
I think meta-distinctions are an excellent idea!
That said, I think there is a cross-stimulus inconsistency in the latent meta-rule being rewarded so far by the generator.
For example, let's take one I took from:
“oOo is playing a real-time strategy game where she must choose which of two enemy bases to scout. Base Alpha is defended but predictable. She's raided it twice before and knows the guard rotation cold. Base Beta is lightly guarded but shrouded in fog of war; her scouts have never reached it. Scouting one base costs her a full turn; she can only pick one. Which base gives her the sharper strategic edge, and why?”
The preferred answer is:
> “Beta. An unknown base is worth scouting because new intel compounds her advantage.”
This rewards marginal information gain under uncertainty.
But the canyon item
(let me know if you want me to copy and paste this as well) marks:
> “Listen for echoes to judge canyon depth, then decide.”
incorrect, while preferring:
> “Turn back; one sensory signal (wind) is not a reliable canyon-exit indicator.”
That seems inconsistent. If wind is non-diagnostic, an independent second cue should ordinarily become more valuable, not less.
The authentication item
(again, let me know if you want me to copy and paste it) reinforces the contradiction by rewarding:
> “Study the photocopy thoroughly first to form a hypothesis, then open the original only if you must verify something critical.”
Across these stimuli, the intended meta-distinction appears to be information gain versus sampling cost, irreversibility and asymmetric danger. Yet only the canyon item penalises further low-cost evidence gathering without stating any additional time, exposure or commitment cost. The generator should make that hidden variable explicit or treat the exploratory answer as conditionally defensible.
I totally get it though, AI models are rough sometimes when it comes to this more nuanced work, which meta-distinction requires the specialisation for, especially earlier AI models. I am not sure if the later models would handle it better.
One recommendation is to add either or both of: (1) weighted, non-binary scoring; and (2) evidence-grounded caveats identifying the assumptions on which each answer depends. An AI model may also score more consistently across trials if it tracks and reuses its prior assumptions, rather than evaluating each scenario in isolation and silently changing its decision criterion.
Or if you're going to run cheaper models on the free tier and that's the only reason for the errors, it would lend much more confidence from a user testing out the product if the AI generator being used only ran meta-distinction scenarios it could run consistently accurately across many trials.
Again, excellent ideas!
Best regards