Qwen-AgentWorld: The Open-Source AI That Simulates Agent Environments
Alibaba's Qwen team has released Qwen-AgentWorld-35B-A3B, the first open-weight Language World Model (LWM) designed to simulate the environments AI agents act in — rather than simply deciding what to do next.
What Is Qwen-AgentWorld?
Most AI research focuses on making agents smarter at choosing actions. Qwen-AgentWorld flips the problem: instead of simply generating responses, it learns to simulate entire environments and predict what will happen after an agent performs an action. In other words, if an agent executes a shell command, clicks a button, or calls an API, Qwen-AgentWorld predicts what the environment would return — essentially acting as a virtual world for AI agents to train and reason in.
Key Technical Details
Model family: Qwen-AgentWorld-35B-A3B (35B total parameters, 3B activated via Mixture-of-Experts) and a larger 397B-A17B variant, both built on the Qwen3.5 base.
Architecture: Uses a hybrid Gated DeltaNet + Gated Attention + MoE layout with a context window of 262,144 tokens — essential for simulating long, multi-turn agent sessions.
License: Apache 2.0 (fully open-weight).
Release date: June 24, 2026, paired with arXiv paper 2606.24597 filed June 23 and model weights published on GitHub.
Seven Unified Domains
Qwen-AgentWorld is the first language world model to cover seven agent interaction domains within a single model: MCP (tool calling), Search, Terminal, SWE (software engineering), Android, Web, and OS. Previous approaches required separate simulators per domain or relied on expensive real-environment infrastructure.
How It Was Trained
The model was trained on more than 10 million environment interaction trajectories across 7 domains through a three-stage pipeline: CPT (Continual Pre-Training) injects general-purpose world modeling capabilities, SFT (Supervised Fine-Tuning) activates next-state-prediction reasoning, and RL (Reinforcement Learning) sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. Crucially, environment modeling is the training objective from the CPT stage onward — making it a native world model, not a general LLM adapted after the fact.
Performance
The model was evaluated on AgentWorldBench, a new benchmark built from real-world interactions of five frontier models across nine established agentic benchmarks. Key results:
Qwen-AgentWorld-35B-A3B scored 56.39 overall, matching Claude Sonnet 4.6 (56.04) and approaching Claude Opus 4.6 (57.80) — despite being a much smaller and open-weight model.
Qwen-AgentWorld-397B-A17B scored 58.71, outperforming GPT-5.4 (58.25) and all other frontier models overall.
The 35B model particularly excels in Search (36.69) and SWE (65.63), closing the gap with models many times its size.
Two Key Use Paradigms
Decoupled Environment Simulator: Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone.
Agent Foundation Model Warm-Up: World-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks.
Why Open-Source Matters Here
Real AI agent training is expensive, fragile, and difficult to scale. Running thousands of browser instances, containers, and virtual machines is expensive. Real-world environments are difficult to duplicate millions of times for reinforcement learning. You cannot easily create rare failure conditions or edge cases on demand. Some environments contain irreversible actions or sensitive data. A freely available, open-weight world model addresses all these pain points.
Why This Is Important
Qwen-AgentWorld points at a problem every serious AI coding team eventually hits: agents do not only need better reasoning models — they need better environments to train against. By releasing a native, open-weight Language World Model that covers seven domains under the permissive Apache 2.0 license, Alibaba/Qwen has made a previously academic concept practically accessible to every developer. This could dramatically reduce the cost and complexity of training capable AI agents, accelerate agentic RL research, and shift the field's focus toward environment modeling as a first-class concern — not an afterthought. If 2025 was the year of coding agents, 2026 is becoming the year of agent environments.
Source here.