Expert verdicts on what AI builds.
Research

The eval space

Models generate games inside our instrumented worlds. Fourteen expert evals answer two questions: did the model build what was asked, and is it any good?

Prompt fidelity
01Game MechanicsDid the model build the mechanic the prompt asked for?

Did the model build the mechanic the prompt asked for? Prompt in, mechanic out, judged against the spec by an expert who plays to probe it.

Rule fidelity Twist / constraint preserved Edge-case behavior Interaction with other systems Exploit & softlock probing Feel matches intent Combat systems Movement Abilities Crafting Progression New features
prompt: "A macro layer where weather changes unit costs" mechanicPresent: true fidelity: 5/7 deviation: "Weather affects build time, not cost" severity: MEDIUM

Judged by: Systems designers, gameplay engineers · Also labeled on: fun factor · learning curve · depth · replayability · frustration

02Spec AdherenceDoes the whole artifact match the prompt?

Does the whole artifact match the prompt? Checked line by line with the prompt in the rater's hand.

Genre match Required mechanics present Constraints respected Content requirements Tone / art direction Scope creep

Judged by: Game designers · scored 1–7 against fixed anchors, per prompt clause

03Iteration AccuracyWhen you prompt a change, does exactly that change?

When you prompt a change, does exactly that change? Experts compare builds before and after and catalog everything else that moved.

Targeted change landed No collateral damage Balance preserved Regression catalog Compounding edits

Matters for: prompt-to-game tools, coding agents, live iteration loops

Design quality
04Level DesignIs the generated environment well designed?
Map flow Navigation Exploration Difficulty curve Player guidance Reward placement Enemy encounters Puzzle quality

Judged by: Level designers, experienced players

05BalanceAre the generated systems balanced?
Characters Weapons Items Cooldowns Progression speed Difficulty spikes Dominant strategies

Judged by: Competitive players, game designers

06NarrativeDoes the generated story work?
Dialogue quality Character consistency Quest structure Pacing Emotional impact Lore consistency

Judged by: Narrative designers, writers

Player experience
07Player ExperienceWill players enjoy this?

Will players enjoy this? The closest category to pure taste.

Fun rating Engagement Frustration Confusion Emotional impact Would you keep playing? Would you recommend?

Note: Our playtest rubric lives here and in design quality

08UX / OnboardingCan players understand and use the game?
Menus HUD Inventory Controls Tutorials Accessibility Information clarity

Judged by: UX designers, onboarding specialists

09AI NPC / Agent InteractionDoes this AI feel intelligent to play with?
Huge future category
NPC conversations Memory consistency Personality Improvisation Realism Safety Player immersion

Judged by: Players in extended sessions, narrative + systems designers

Generated artifacts
10Generated ContentIs AI-generated content good?

Is AI-generated content good? Humans rank quality, creativity, usefulness, and consistency.

Quests Characters Weapons Levels Dialogue Textures Animations

Note: Whole generated games, the top of this stack, are where we start

11Code / TechnicalIs AI-generated game code good?
Unity scripts Unreal Blueprints Shaders Gameplay systems Networking Optimization

Judged by: Unity/Unreal developers · Labels: does it work? performant? maintainable? would you ship it?

12QA / StabilityDoes the generated build hold up?

Does the generated build hold up? Every crash, softlock, and exploit documented with reproduction steps.

Crash / softlock detection Severity ranking Reproduction steps Regression testing Edge cases

Judged by: QA engineers

Systems & economy
13Game EconomyIs the generated economy healthy?

Is the generated economy healthy? Experts stress the loop for inflation, dead ends, and degenerate strategies.

Currency inflation Progression pacing Reward systems Retention loops Sink / source balance

Judged by: Economy designers, live-ops specialists

Preference data
14Human PreferencesA or B: which is better, why, and by how much?
The universal one

A or B: which is better, why, and by how much? An expert picks the winner, states a confidence, and writes the reason.

Pairwise choice Written rationale Confidence: slight / clear / strong

Used for: RLHF · reward models · benchmarking

Fourteen questions. One harness.