Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning | Summary
27 Aug 2026 | Paper Review Physical Reasoning Executable World Representations Agentic Discovery Vision-Language ModelsContents
- Summary
- 2 In Search of Physical World Representations
- 3 Code as Worlds: Executable World Representations
- 4 Agentic Discovery of World Representations
- 5 Application: Learning Quantitative Physical Reasoning
- Appendix
- Brief Thoughts
This article explains the key points of Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning.
- 2026-08-27 (arXiv), MirroS Technical Report
- Wang, Hanyang, Cai, Yimo, Chen, Weiliang, Chi, Jiawei, Sun, Haowen, Dai, Qiyu, Hung, Yi-Hsin, Guo, Xingzhuo, Ren, Jinshan, Yao, Runmao, et al.
- MirroS, Tsinghua University, Peking University, Nanyang Technological University
- Paper
- Github
- Project Page
Summary
- Code-as-World represents a physical scene as executable code rather than only pixels, latent features, or language, making objects, states, parameters, dynamics, and interventions explicit. It constructs these representations from text or video with an iterative simulation-and-verification agent, then uses verified worlds as scalable supervision for quantitative physical reasoning in vision-language models. On QuantiPhy-validation, Code-as-World-VL-27B (Reasoning) reports 58.6 macro MRA, while the controlled 4B and 9B direct-answer variants report 50.6 and 55.4.
Introduces the proposed interface: code specifies composition, evolution, and appearance, which the paper presents as support for physical questions, controllable video generation, and embodied-action planning.
2 In Search of Physical World Representations
The paper argues that pixel prediction can yield visually plausible futures without identifying their physical causes, while 3D reconstruction can preserve geometry without disentangling causal dynamics. Language is compact and compositional but imprecise for continuous quantities such as trajectories, contacts, and physical parameters; the proposed representation is intended to combine semantic abstraction, structure, and temporal mechanisms.
3 Code as Worlds: Executable World Representations
An executable world representation (EWR) is defined as p = (C, E, A) ∈ Pexec. C specifies physical composition—objects, geometry, dimensions, mass, friction, gravity, and static supports—E specifies initial states, temporal evolution, events, and duration, and A specifies camera, materials, lighting, backgrounds, resolution, and rendering settings. The implementation couples this inspectable and editable representation to a simulation engine, prioritizing physical equivalence over pixel-level duplication.
Contrasts pixel/video, 3D, and language representations with executable code. The figure frames Code-as-World as adding explicit continuous physical mechanisms to visual detail, geometric structure, and compact semantics.
4 Agentic Discovery of World Representations
World construction is formulated as abductive search over executable hypotheses rather than one-shot prediction. Given modality-specific evidence, the agent proposes an EWR, instantiates simulator-ready parameters, executes a trajectory, renders predicted observations, verifies discrepancies, and revises the relevant world component until acceptance or exhaustion of the iteration budget.
4.1 Evidence Construction
For text inputs, the system extracts entities, spatial relations, events, and intended outcomes, then fills unspecified geometry, parameters, and camera settings with priors and defaults. For real videos, it uses depth maps, instance masks, tracks, camera geometry, and generated object meshes to constrain object scale, position, and dynamics; candidate worlds are retained only after rendering and verification against the source video.
4.2 Agentic Discovery Loop
At each round, the agent updates p using evidence η and structured feedback Δ, compiles it into simulator parameters θ, and executes a state trajectory τ containing states, contacts, collisions, and outcomes. For video, verification compares rendered RGB, projected depth, masks, and image-plane tracks at selected frames; for text, it focuses on semantic and physical constraints. The procedure rejects a hypothesis if it does not satisfy the acceptance condition within K iterations.
Shows the shared discovery loop for text and video inputs. Semantic evidence from language or RGB, depth, instance, and track evidence from video constrains an EWR that is proposed, instantiated, executed, rendered, compared, and refined.
4.3 Evaluation
For video-driven reconstruction, the authors filter WISA-80K for salient rigid-body motion, excluding clips with substantial camera motion, insufficient object motion, severe editing, or incomplete events, then manually review the remainder. They evaluate up to K = 5 rounds with Visual Alignment, Object IoU, Traj-ADE, Velocity-ADE, and Accuracy@2%D; at an equal five-evaluation budget, the main experiment exceeds Best-of-5 on Visual Alignment, Object IoU, Traj-ADE, and Accuracy@2%D. The appendix reports that, with the physics engine, the fifth round exceeds the matched Best-of-5 baseline on all five metrics.
Demonstrates editability after world construction: changing bowling directions or camera viewpoints changes the simulator rollout and its aligned realistic-video rendering.
Plots reconstruction metrics across one-shot generation and five iterative rounds. At the matched five-evaluation budget, the final iterative result exceeds the Best-of-5 reference for Visual Alignment, Object IoU, Traj-ADE, and Accuracy@2%D, but not Velocity-ADE.
4.4 Limitations
The method is limited by simulator coverage: small unobserved differences in terrain, contact geometry, material properties, or other latent conditions can alter rigid-body motion. The loop can therefore converge to an EWR that is locally plausible and visually aligned without providing a mechanistically correct account when the underlying process lies outside the simulator’s modeling scope.
Provides text-to-world examples in which simulation rollouts generated from prompts are transformed into more realistic videos while the paper claims that their encoded trajectories and physical evolution are preserved.
Provides real-video-to-simulation examples. The abstracted simulations reproduce selected scenes’ coarse composition, motion, and interactions rather than their full photorealistic appearance.
5 Application: Learning Quantitative Physical Reasoning
The downstream task predicts a scalar physical quantity from a monocular video and question, covering size, displacement, velocity, and acceleration. World-space questions supply a known reference quantity ρ; following QuantiPhy, the target is calibrated from pixel-space measurement using γ = ρ/ρpix and y = γypix. The paper trains 4B and 9B direct-answer variants with image-space supervision followed by executable-world supervision, and separately trains a 27B reasoning variant.
Illustrates QuantiPhy’s four settings: 2D or depth-aware 3D scenes, each paired with a static or dynamic prior used to calibrate a requested world-space measurement.
5.2 Image-Space Measurement Grounding
Image-space supervision is derived from bounding boxes, masks, and tracks in RefCOCO, RefCOCO+, RefCOCOg, RefCLEF, and GOT-10K. Object size and position come from boxes or masks, while motion comes from tracked centers using d = ct2 − ct1, vt = (ct+1 − ct−1)/(2Δt), and at = (ct+1 − 2ct + ct−1)/Δt²; supervised fine-tuning on Dpix establishes localization, measurement, and tracking before metric calibration.
5.3 World-Space Physical Calibration from Verified Executable Worlds
Verified EWRs provide synchronized videos and exact scene geometry, camera settings, timestamps, and physical-state trajectories, enabling world-space VQA labels from Dtext and Dvideo. The second phase uses GRPO with numerical, unit, and format rewards; text-driven worlds supply exact simulator labels, whereas video-driven worlds better match real visual and motion distributions. On QuantiPhy-validation, the 9B direct-answer model reports 55.4 average MRA versus 40.2 for Qwen3-VL-32B-Instruct and 54.8 for Gemini-3.1 Flash; the 27B reasoning model reports 58.6, but its scale and response protocol differ from the direct-answer comparison.
Reports QuantiPhy-validation MRA by 2S, 2D, 3S, and 3D categories. Code-as-World-VL-4B, 9B, and 27B (Reasoning) score 50.6, 55.4, and 58.6 average MRA, respectively; the 27B result is scored from the answer emitted after its reasoning trace.
Places QuantiPhy average MRA against model parameter count. The plotted Code-as-World series rises from 50.6 at 4B to 55.4 at 9B and 58.6 for the 27B reasoning model, while models with undisclosed sizes are grouped separately.
Compares simulator renders with sim-to-real videos. Sim-to-real improves distributional realism from JEDi MMD 3.000 to 1.484 and TRAJAN Fréchet 406.872 to 185.321, while Traj-ADE changes from 1.682 to 1.677 %D, Velocity-ADE worsens from 0.404 to 0.472 %D/step, and Accuracy@2%D falls from 78.81% to 77.49%.
Repeats iterative discovery with the physics engine rather than the animation engine. Across five rounds, Visual Alignment, Object IoU, and Accuracy@2%D increase while Traj-ADE and Velocity-ADE decrease; the final result exceeds matched Best-of-5 on all five metrics.
Reports pixel-level MRA across four referring-expression datasets and GOT-10K. Adding executable-world reinforcement learning improves both sizes on every listed benchmark; for example, the 9B model rises from 63.7 to 68.3 on RefCOCO and from 20.1 to 26.6 on GOT-10K.
Qualitatively compares grounding, length, speed, and acceleration predictions. In the selected examples, the Code-as-World-VL outputs are closer to the displayed target boxes and numerical ground truth than the shown comparison outputs.
Separates the contributions of image-space data Dpix and text- and video-driven world data Dtext and Dvideo. Each world-data source improves the listed average over image-space-only training, and using both sources gives the strongest tabulated averages: 50.6 for 4B and 55.4 for 9B.
Displays a video-driven soccer-ball reconstruction with portions of scene.json and simulation_sdk.py. The color coding maps structured fields and SDK operations to composition, evolution, and appearance, and the implementation executes the scene in MuJoCo.
Appendix
- The appendix specifies 73,335 Image-Space question-answer pairs and an executable-world dataset with 1,585 text-driven and 988 video-driven VQA samples. It details MuJoCo-based animation and physics engines, QuantiPhy’s 159 validation pairs, MRA evaluation, reasoning-model settings, realism and motion metrics, and data-source ablations. The ablation table shows that combining Dpix, Dtext, and Dvideo yields the strongest listed QuantiPhy averages: 50.6 for 4B and 55.4 for 9B.
Brief Thoughts
The central contribution is not only simulation-data generation: it makes physical assumptions inspectable, editable, and testable through an executable representation. The reported evidence supports gains on constrained rigid-body reconstruction and QuantiPhy calibration, but broader claims about physical intelligence remain contingent on simulator fidelity, coverage of broader physical regimes, camera-motion robustness, and learning the discovery loop itself rather than only from its verified outputs.