World Editing: Intervening on Executable Worlds at Increasing Depth
Abstract
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
Community
Interactive world models are increasingly able to generate environments and act within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on a world while preserving the properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples a world’s entities, dynamics, and systems. We instantiate the capability through industry-grade game modding: IGMWorld provides the executable environment, and IGMBench contributes 110 tasks with over 1.1K deterministic state and behavioral criteria across Minecraft and Terraria.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development (2026)
- GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions (2026)
- GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? (2026)
- SWE-Game: Can Coding Agents Build the Games We Want? (2026)
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories? (2026)
- GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents (2026)
- GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02331 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper

