Agentic verification · Progressive development

RSIGame

Autonomous Agentic Game Development with Recursive Self-improvement

Wenyi Wu1*, Minghao Fu1,2*, Jieyu You2, Kun Zhou1†, Siqi Liu1, Aayush Salvi1, Yiheng Lin2, Ce Zhang3, Xiaohan Lan2, Jiahui Zhu2, Yujie Zhong2†, Qi She2, Biwei Huang1

1University of California San Diego   2ByteDance Inc.   3Carnegie Mellon University

*Equal contribution   †Corresponding authors

+14.3Overall on Codex + GPT-5.5 · 50.3 → 64.5
61.4Qwen3.8-27B games after experience internalization and RSIGame development — past one-shot GPT-5.5 (50.3)
1,998runnable agent-written games, to be released
222Kverified agentic training rows, to be released
RSIGame in 78 seconds: the local and global loops at work, one game developed across stages with sparse high-level guidance, and before/after comparisons across 2D and 3D games.

Abstract

Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions.

We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training.

Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.

Before / After

The ten largest gains we have on record, five per generator. Each pair replays the same scripted inputs on two builds of the same game: the initial generation (round 0) and the build development delivers. Sweep across the frame to wipe between them; pick another game from the ten below.

Play the before and after

Six games taken all the way through development, in the browser. Switch between the initial generation and the delivered build — the same game, the same controls, whatever the rounds changed. Every intermediate round lives in that game's own Space.

How far 30 rounds go

GameCraft-Bench scores a game by replaying scripted demos and judging them against a per-task rubric: Overall = Build × (0.15·Mechanics + 0.35·Depth + 0.15·Visuals + 0.35·Art), shown ×100. Every method starts from the same frozen base P₀ and is scored by the same judge on the same demo scripts.

Development-time scaling

Mean Overall of the build a budget of k rounds delivers · band = ±1 s.e. over games · hover for values

A free-form playtest-and-revise baseline finishes a strong base where it started; the Global Quality Monitor is what keeps late rounds from undoing earlier gains. Round 30 minus round 0 under each panel.

Rubric categories

All 140 Godot tasks · frozen base (hollow) → RSIGame (filled)

Art and Depth carry the most weight in Overall (0.35 each) and are where the frozen bases are weakest — GPT-5.5 starts at 44.6 Art, Qwen3.8-27B at 33.4 Depth.

Main results, all 140 tasks

Citation

@misc{wu2026rsigame,
  title  = {RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement},
  author = {Wu, Wenyi and Fu, Minghao and You, Jieyu and Zhou, Kun and Liu, Siqi and
            Salvi, Aayush and Lin, Yiheng and Zhang, Ce and Lan, Xiaohan and Zhu, Jiahui and
            Zhong, Yujie and She, Qi and Huang, Biwei},
  year   = {2026},
  eprint = {2609.39045},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url    = {https://arxiv.org/abs/2609.39045}
}