
Voyager (2023) solved agentic tasks with code execution
Beating Minecraft with Code Execution
In mid-2023, a research project called Voyager made waves: it effectively solved Minecraft, performing several multiples better than the prior SOTA. This was a massive breakthrough as previous reinforcement learning systems had struggled for years with even basic Minecraft tasks. While the AI community was focused on scaling intelligence, Voyager demonstrated something more fundamental: the right tools can unlock entirely new tiers of capability. The same GPT-4 model that struggled with Minecraft using standard agentic frameworks (like ReAct) achieved remarkable results when allowed to write and execute code. This wasn’t about raw intelligence—it was about giving the agent a more expressive way to act.
equipItem(...) (this would be typical of “traditional” agent algorithms, such as ReAct), it could create higher-level operations like craftShieldWithFurnace() through composing the atomic APIs. Furthermore, Wang et al. implemented a memory mechanism, in which these successful “action programs” could later be recalled, copied, and built upon, effectively enabling the agent to accumulate experience.

Code is an Ideal Action Space
What these authors demonstrated is a fundamental insight that extends far beyond gaming. Letting AI act through code rather than atomic commands will lead to a step change in the capabilities of AI systems. Nowhere is this more apparent than in software engineering, where agents already understand complex transformations but lack the tools to execute them effectively. Today’s productionized code assistants operate though an interface where they can directly read/write to text files and perform other bespoke activities, like searching through file embeddings or running terminal commands. In the act via code paradigm, all of these actions are expressed through writing and executing code, like the below:- API-Driven Extensibility: Any operation that can be expressed through an API becomes accessible to the agent. This means the scope of tasks an agent can handle grows with our ability to create clean APIs for complex operations.
- Programmatic Efficiency: Many agent tasks involve systematic operations across large codebases. Expressing these as programs rather than individual commands dramatically reduces computational overhead and allows for batch operations.
- Composability: Agents can build their own tools by combining simpler operations. This aligns perfectly with LLMs’ demonstrated ability to compose and interpolate between examples to create novel solutions.
- Constrained Action Space: Well-designed APIs act as guardrails, making invalid operations impossible to express. The type system becomes a powerful tool for preventing entire classes of errors before they happen.
- Objective Feedback: Code execution provides immediate, unambiguous feedback through stack traces and error messages—not just confidence scores. This concrete error signal is invaluable for learning.
- Natural Collaboration: Programs are a shared language between humans and agents. Code explicitly encodes reasoning in a reviewable format, making actions transparent, debuggable, and easily re-runnable.