Models generate games inside our instrumented worlds. Fourteen expert evals answer two questions: did the model build what was asked, and is it any good?
Did the model build the mechanic the prompt asked for? Prompt in, mechanic out, judged against the spec by an expert who plays to probe it.
Does the whole artifact match the prompt? Checked line by line with the prompt in the rater's hand.
When you prompt a change, does exactly that change? Experts compare builds before and after and catalog everything else that moved.
Will players enjoy this? The closest category to pure taste.
Is AI-generated content good? Humans rank quality, creativity, usefulness, and consistency.
Does the generated build hold up? Every crash, softlock, and exploit documented with reproduction steps.
Is the generated economy healthy? Experts stress the loop for inflation, dead ends, and degenerate strategies.
A or B: which is better, why, and by how much? An expert picks the winner, states a confidence, and writes the reason.