Agentic verification · Progressive development
Autonomous Agentic Game Development with Recursive Self-improvement
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions.
We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training.
Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
The ten largest gains we have on record, five per generator. Each pair replays the same scripted inputs on two builds of the same game: the initial generation (round 0) and the build development delivers. Sweep across the frame to wipe between them; pick another game from the ten below.
Six games taken all the way through development, in the browser. Switch between the initial generation and the delivered build — the same game, the same controls, whatever the rounds changed. Every intermediate round lives in that game's own Space.
The GameCraft-Bench games with a replayable before/after pair, from GPT-5.5 and Qwen3.8-27B initial generations in Godot 4. Runs whose retained checkpoint is round 0, or whose two builds share no replayable demo, are not shown; the Results section reports all 140 tasks. Open a game for the side-by-side replay, the score of every checkpoint, what each of the 30 rounds did, and the judge's reasoning before and after.
GameCraft-Bench scores a game by replaying scripted demos and judging them against a per-task rubric: Overall = Build × (0.15·Mechanics + 0.35·Depth + 0.15·Visuals + 0.35·Art), shown ×100. Every method starts from the same frozen base P₀ and is scored by the same judge on the same demo scripts.
Mean Overall of the build a budget of k rounds delivers · band = ±1 s.e. over games · hover for values
A free-form playtest-and-revise baseline finishes a strong base where it started; the Global Quality Monitor is what keeps late rounds from undoing earlier gains. Round 30 minus round 0 under each panel.
All 140 Godot tasks · frozen base (hollow) → RSIGame (filled)
Art and Depth carry the most weight in Overall (0.35 each) and are where the frozen bases are weakest — GPT-5.5 starts at 44.6 Art, Qwen3.8-27B at 33.4 Depth.
@misc{wu2026rsigame,
title = {RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement},
author = {Wu, Wenyi and Fu, Minghao and You, Jieyu and Zhou, Kun and Liu, Siqi and
Salvi, Aayush and Lin, Yiheng and Zhang, Ce and Lan, Xiaohan and Zhu, Jiahui and
Zhong, Yujie and She, Qi and Huang, Biwei},
year = {2026},
eprint = {2609.39045},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.39045}
}