World Editing: Intervening on Executable Worlds at Increasing Depth

Max Ku1,2, Nok-Kan Law2, Yu-Chien Tang2, Shih-Ying Yeh2,3, Ping Nie1, Andy Zheng1, Tat Hei Lai1, Fei-Yueh Chen2, Nikko Yu2, Wei-Chieh Sun2, Suzy Huang2, Chiao-Wei Hsu2, Chih-Chuan Huang2, Chak-Wing Mak2, Ho Yin Sam Ng2, Edisy Kin Wai Chan2, Min-Hung Chen2, Ho Kei Cheng2,4
1University of Waterloo, 2G-G-G, 3Comfy Org Research, 4University of Illinois Urbana-Champaign Contact: m3ku@uwaterloo.ca
Successful world edits in Minecraft and Terraria across intervention levels
Agents intervene on commercially released games through their native modding ecosystems, from introducing a single new entity to rebuilding a coupled world-level system. Top: Minecraft. Bottom: Terraria.

TL;DR

Interactive world models are increasingly able to generate environments and act within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on a world while preserving the properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples a world’s entities, dynamics, and systems. We instantiate the capability through industry-grade game modding: IGMWorld provides the executable environment, and IGMBench contributes 110 tasks with over 1.1K deterministic state and behavioral criteria across Minecraft and Terraria.

  • Frontier agents are already capable, but reliability falls with depth. The strongest configuration solves 78.2% of tasks under a strict all-criteria measure. Every agent peaks on shallow property edits and degrades on deeper dynamics and system edits, and the effect survives controlling for how many criteria a task has, so depth is not simply “more requirements.”
  • Failures happen after the build succeeds. Most unsuccessful edits still compile and load. The bottleneck is realizing the requested behavior, and the failures that grow with depth are the ones where new content misbehaves on contact with existing systems.
  • Visual consistency is a separate axis. Native game art reaches a joint style-and-semantic pass rate of 0.80 in Minecraft and 0.78 in Terraria; every evaluated agent stays below 0.50. The visual ranking does not track the functional one, so world editing is multidimensional.
  • Evaluation is deterministic. Every criterion is an executable check against the running game: state queries, scripted in-game actions, and paired regression checks, with no LLM or VLM acting as judge.
110
World-editing tasks
1,161
State & behavioral criteria
7
Agent configurations
78.2%
GPT-5.6 Sol, overall
48.1%
GPT-5.6 Sol at L4
<0.50
Joint visual pass, all agents

World Editing as Intervention

We formulate world editing as an intervention on an existing executable world. Let a world be abstractly represented as \(W = (S, T, R, \dots)\), where \(S\) denotes its entities and state space, \(T\) its transition dynamics, and \(R\) the rules and systems that govern interactions. Given a natural-language editing instruction, a world-editing agent transforms \(W\) as

$$W \xrightarrow{\;\Delta_k\;} W' = f(W, \Delta_k), \qquad k \in \{\text{property},\, \text{entity},\, \text{dynamics},\, \text{system}\}$$

where \(\Delta_k\) denotes an intervention of type \(k\), and \(W'\) is the resulting world that realizes the requested intervention while preserving unrelated properties of \(W\). We characterize interventions by their depth, running from property edits (L1) through entity and dynamics changes to system-level interventions (L4). In game-modding terminology these correspond to parameter, content, mechanic, and system editing.

The levels characterize the semantic scope and coupling of a requested intervention rather than prescribing which implementation components must change. Intervention depth therefore measures how strongly an edit must coordinate multiple parts of the world, not how much code it requires: a system-level edit may be small in implementation yet still need several interacting components to remain consistent.

Level World Intervention Game Modding Examples
L1 Property Intervention: Modify properties of existing components without changing their identity or underlying behavior. Parameter Editing: Modify parameters, configurations, or appearance of existing game components. Increase weapon damage; replace a monster's visual asset.
L2 Entity Intervention: Introduce new entities or content while largely preserving existing dynamics. Content Editing: Add new content that conforms to existing mechanics. Add a crafting recipe or item; add a monster or NPC.
L3 Dynamics Intervention: Modify rules governing entity behavior and interaction. Mechanic Editing: Modify game logic governing behaviors and interactions. Add an air-dash mechanic; add a status-effect system.
L4 System Intervention: Modify multiple coupled components and interactions forming a world-level system. System Editing: Modify multiple interacting game subsystems. Add a skill tree; overhaul the economy system.
World-editing taxonomy and its realization in game modding.

IGMWorld and IGMBench

Two components support this evaluation. IGMWorld provides the executable environment in which agents modify existing game repositories, build and run their edits, and are evaluated on the resulting world rather than on code structure alone. IGMBench provides the task instances: 110 world-editing tasks across Minecraft and Terraria, distributed approximately evenly across the four intervention levels, with 1,161 state and behavioral criteria in total.

Overview of IGMWorld
Overview of IGMWorld. A world-editing agent modifies a scaffold mod repository from a natural-language request and produces an edited world, which is evaluated for executability, state and behavioral correctness, and visual consistency with existing in-game assets.

How Edits Are Evaluated

Every one of the 1,161 criteria is a deterministic check against the running game. No language or vision–language model is asked to judge whether an edit succeeded. Validators assess observable world behavior rather than source-code structure, so different implementations may satisfy the same intervention, and a task succeeds only if it builds and passes all of its criteria.

Stage 1
Executability
The submission must compile into a valid mod package and load in the target game without crashing. Build and load are strict gates.
Stage 2
World-Edit Correctness
State checks query properties of the running game: entity attributes, item statistics, recipes, registrations. Behavioral checks execute controlled in-game actions and verify the resulting state transitions. Paired regression checks (225 across 30 tasks) compare the modified and original worlds to test that unrelated properties remain unchanged.
Stage 3
Visual Consistency
An asset is scored by how close it sits to the game’s own art: the mean distance to its \(k=5\) nearest category-matched vanilla references (TPIPS supplies the text-conditioned distance). The pass bar is not chosen by hand. It is the \(\alpha=0.15\) quantile of leave-one-out scores among the vanilla assets themselves, so native art passes at 0.85 by construction.

The reference library is the game’s own art: 1,989 assets in Minecraft and 15,456 in Terraria. The same construction runs under two text factors, the requested object identity for semantic consistency and “art style” for style consistency, so “consistent” means as close to the game’s art as the game’s art is to itself. We checked the choices for stability (k from 1 to 10 preserves at least 88% of pass/fail verdicts) and against people: nine raters with prior gaming experience judged 40 blind-sampled agent textures, and the style factor agrees with the majority human judgment at 0.66 AUC, against 0.54 for a CSD style embedding.

$$S_f(x; c) \;=\; \frac{1}{k} \sum_{r_i \,\in\, \mathrm{kNN}(x,\, I_c;\, f)} \mathrm{TPIPS}(x, r_i; f)$$

Here \(x\) is the evaluated asset, \(I_c\) the native reference assets of category \(c\), and \(f\) the text factor. An asset passes when \(S_f(x;c) \ge \tau_{c,f}\), where the threshold \(\tau_{c,f}\) is the \(\alpha\)-quantile of leave-one-out scores over \(I_c\).

Reliability Varies With Intervention Depth

Current frontier agents already exhibit substantial functional world-editing capability. Criterion Pass Rate runs well above the stricter World-Editing Success Rate throughout, indicating that many unsuccessful tasks nevertheless satisfy substantial portions of the requested behavior.

Reliability varies systematically with intervention depth. Every evaluated configuration achieves its highest WSR at L1, while performance is generally lower for deeper dynamics and system interventions, although the relative ordering of L3 and L4 varies across models and games. Because deeper tasks also contain more evaluation criteria, we stratify tasks by criterion count: within comparable criterion-count ranges, L1 remains substantially more reliable, suggesting that intervention depth captures structure beyond criterion count alone.

World-editing success rate by intervention level
(a) WSR by intervention level
World-editing success rate by criterion-count bin
(b) WSR controlled by criterion count
Reliability across intervention depth. Performance differences persist within comparable criterion-count ranges, showing that intervention depth is not reducible to requirement count alone.
ModelL1L2L3L4Overall
CPRWSRCPRWSRCPRWSRCPRWSRCPRWSR
DeepSeek-V4-Pro92.174.180.824.147.718.547.011.163.331.8
GLM-5.397.692.682.762.139.825.939.311.160.148.2
Kimi-K398.588.978.365.561.051.973.333.374.960.0
Gemini 3.5 Flash87.777.886.455.283.537.068.859.381.057.3
Claude Opus 4.899.692.692.872.494.766.789.455.693.671.8
GPT-5.6 Luna96.392.690.258.690.444.480.063.088.564.5
GPT-5.6 Sol99.896.393.986.298.081.588.848.194.878.2
Criterion Pass Rate (CPR) and World-Editing Success Rate (WSR) on IGMBench, broken down by world-intervention level. CPR measures the fraction of individual state and behavioral criteria satisfied; WSR measures the fraction of tasks for which all associated criteria are satisfied. Results are aggregated across Minecraft and Terraria; visual correctness is evaluated separately below. Bold: best; underline: second best.

Most Failures Occur After Executability

Most unsuccessful world edits fail after reaching an executable state. Behavioral failures account for the large majority of unsuccessful runs at every intervention level, while build and load failures remain comparatively uncommon. Even at L4, where build failures become more frequent, behavioral correctness remains the dominant failure stage. The principal bottleneck is therefore not producing a modification that compiles and loads, but realizing the requested behavior in the executable world.

The composition of behavioral failures also changes with depth. L1 failures are almost entirely mechanic semantics, such as a wrong value or formula. From L2 onward, interaction/progression failures, in which new content misbehaves on contact with existing systems, become the largest category and keep growing with depth, while resource/registration failures stay flat. This is consistent with deeper edits requiring coordination across multiple interacting world components rather than a single localized change.

Failure stage by intervention level
(a) Failure stage by intervention level
Behavioral failure domain by intervention level
(b) Behavioral failure domain by level
Failure structure across intervention depth. Behavioral failures dominate across all levels. Among behavioral failures, interaction/progression failures increase from L2 to L4, while resource/registration failures remain relatively stable. Numbers inside bars denote failure counts.
Representative failure cases across intervention levels
Representative failure cases across intervention levels in Minecraft (top) and Terraria (bottom). Columns show entity (L2), dynamics (L3), and system (L4) interventions. Failures include missing registration or assets, incorrect mechanic execution, and unintended system-level behavior.

Visual Consistency Is a Distinct Bottleneck

Visual integration remains substantially weaker than functional world editing. In both games, native assets clear the joint style-and-semantic check at a rate no evaluated agent approaches.

Visual performance does not follow the ranking observed for functional world editing. Configurations that are weaker on executable state and behavioral correctness can nevertheless perform competitively on visual assets, while the strongest functional agents do not dominate the visual evaluation. This weak correspondence suggests that world editing is multidimensional: successfully modifying the behavior of an executable world does not imply that newly introduced content integrates perceptually with that world.

MinecraftTerraria
Model / SourceNPassstylePasssemPassbothNPassstylePasssemPassboth
Vanilla (LOO)1,9890.850.850.8015,4560.850.850.78
DeepSeek-V4-Pro1420.170.230.151130.160.230.11
GLM-5.32260.490.530.40730.320.370.29
Kimi-K32050.520.540.471090.290.350.24
Gemini 3.5 Flash1100.320.360.221040.320.320.22
Claude Opus 4.81270.340.420.231100.350.340.25
GPT-5.6 Luna1580.160.200.141100.180.360.12
GPT-5.6 Sol1550.230.240.19990.220.370.18
Visual quality of agent-produced assets. N counts registered texture assets; vanilla textures and missing files fail all checks. Passstyle, Passsem, and Passboth denote style, semantic, and joint pass rates. Vanilla (LOO) provides the in-game reference, with marginal pass rates of 1−α=0.85 by construction. Bold: best agent; underline: second best.
Minecraft item textures grouped by generation method
(a) Minecraft
Terraria item textures grouped by generation method
(b) Terraria
Item textures grouped by generation method. In each panel: left, image-generated; middle, native game assets; right, programmatically drawn. Minecraft image-generated assets tend toward smoother shading and higher contrast, while programmatic assets are flatter and more geometric than native textures. In Terraria, image-generated assets more closely reproduce the game’s dense and outlined visual style, while programmatic assets remain visibly simpler. Each Minecraft panel shows 36 sampled item textures; Terraria panels sample from the same size range.

Beyond Minecraft and Terraria

Although IGMBench quantitatively evaluates Minecraft and Terraria, the formulation is not specific to these games. World editing is better viewed as a general capability over executable worlds, with games serving as a concrete and verifiable testbed rather than the boundary of the problem.

World-editing examples in PEAK, Palworld, and Starbound
World editing beyond the games included in IGMBench. Qualitative examples of property, entity, and dynamics interventions in PEAK, Palworld, and Starbound. These illustrate that the intervention taxonomy applies to additional executable game worlds, but they are not part of the benchmark.

Citation

Citation will be available soon.