Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate

📰 ArXiv cs.AI

Learn to evaluate LLM-generated executable game scenes using a multi-axis approach beyond compile-pass rate, improving the reliability of AI-generated code

advanced Published 11 May 2026
Action Steps
  1. Apply the Mage evaluation protocol to LLM-generated code
  2. Configure a four-axis evaluation system for compile success, runtime success, structural fidelity, and mechanism adherence
  3. Run experiments to test the effectiveness of the Mage protocol on multiple LLMs
  4. Compare the results of different LLMs using the Mage protocol
  5. Test the generated code for runtime success and structural fidelity
Who Needs to Know This

AI engineers, game developers, and researchers can benefit from this approach to improve the evaluation of LLM-generated code, ensuring it meets the required standards

Key Insight

💡 A single metric like compile-pass rate is not enough to evaluate the quality of LLM-generated code, a multi-axis approach is necessary

Share This
🚀 Evaluate LLM-generated game scenes with Mage, a multi-axis protocol that goes beyond compile-pass rate 🚀

Key Takeaways

Learn to evaluate LLM-generated executable game scenes using a multi-axis approach beyond compile-pass rate, improving the reliability of AI-generated code

Full Article

Title: Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate

Abstract:
arXiv:2605.07342v1 Announce Type: cross Abstract: Compile-pass rate is the dominant evaluation signal for LLM code generation, yet for multi-component domain-specific artifacts it can be actively misleading. We demonstrate this on executable game scene synthesis with a four-axis evaluation protocol (named `Mage') -- compile success, runtime success, structural fidelity, and mechanism adherence -- applied to 858 generation attempts across four open-weight LLMs (7B--30B), 26~hand-crafted Unity goa
Read full paper → ← Back to Reads

Related Videos

5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Kimi K3: The Free AI That Just Beat Claude at Coding (Ranked #1)
Kimi K3: The Free AI That Just Beat Claude at Coding (Ranked #1)
AI Andy
GLM-5.2 Is INSANE – Is it The BEST New Open Source Model?
GLM-5.2 Is INSANE – Is it The BEST New Open Source Model?
AI Andy
I Gave Fable 5 Six Impossible Prompts (One Shot Each)
I Gave Fable 5 Six Impossible Prompts (One Shot Each)
AI Andy
EVERY Loop From Matthew Berman's New Loop Library! (Copy & Paste!)
EVERY Loop From Matthew Berman's New Loop Library! (Copy & Paste!)
AI Andy
Ollama + OpenWebUI: Run LLM's Locally For FREE!!
Ollama + OpenWebUI: Run LLM's Locally For FREE!!
Thomas Janssen