Gorio Tech Blog search

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning | Summary

|

Contents

This article explains the key points of Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning.

  • 2026-08-27 (arXiv), MirroS Technical Report
  • Wang, Hanyang, Cai, Yimo, Chen, Weiliang, Chi, Jiawei, Sun, Haowen, Dai, Qiyu, Hung, Yi-Hsin, Guo, Xingzhuo, Ren, Jinshan, Yao, Runmao, et al.
  • MirroS, Tsinghua University, Peking University, Nanyang Technological University
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • Code-as-World represents a physical scene as executable code rather than only pixels, latent features, or language, making objects, states, parameters, dynamics, and interventions explicit. It constructs these representations from text or video with an iterative simulation-and-verification agent, then uses verified worlds as scalable supervision for quantitative physical reasoning in vision-language models. On QuantiPhy-validation, Code-as-World-VL-27B (Reasoning) reports 58.6 macro MRA, while the controlled 4B and 9B direct-answer variants report 50.6 and 55.4.

Introduces the proposed interface: code specifies composition, evolution, and appearance, which the paper presents as support for physical questions, controllable video generation, and embodied-action planning.

Code-as-World represents physical worlds as executable code for physical intelligence.
Code-as-World represents physical worlds as executable code for physical intelligence.

2 In Search of Physical World Representations

The paper argues that pixel prediction can yield visually plausible futures without identifying their physical causes, while 3D reconstruction can preserve geometry without disentangling causal dynamics. Language is compact and compositional but imprecise for continuous quantities such as trajectories, contacts, and physical parameters; the proposed representation is intended to combine semantic abstraction, structure, and temporal mechanisms.

3 Code as Worlds: Executable World Representations

An executable world representation (EWR) is defined as p = (C, E, A) ∈ Pexec. C specifies physical composition—objects, geometry, dimensions, mass, friction, gravity, and static supports—E specifies initial states, temporal evolution, events, and duration, and A specifies camera, materials, lighting, backgrounds, resolution, and rendering settings. The implementation couples this inspectable and editable representation to a simulation engine, prioritizing physical equivalence over pixel-level duplication.

Contrasts pixel/video, 3D, and language representations with executable code. The figure frames Code-as-World as adding explicit continuous physical mechanisms to visual detail, geometric structure, and compact semantics.

Comparison of pixel, 3D, language, and executable-code representations for physical worlds.
Comparison of pixel, 3D, language, and executable-code representations for physical worlds.

4 Agentic Discovery of World Representations

World construction is formulated as abductive search over executable hypotheses rather than one-shot prediction. Given modality-specific evidence, the agent proposes an EWR, instantiates simulator-ready parameters, executes a trajectory, renders predicted observations, verifies discrepancies, and revises the relevant world component until acceptance or exhaustion of the iteration budget.

4.1 Evidence Construction

For text inputs, the system extracts entities, spatial relations, events, and intended outcomes, then fills unspecified geometry, parameters, and camera settings with priors and defaults. For real videos, it uses depth maps, instance masks, tracks, camera geometry, and generated object meshes to constrain object scale, position, and dynamics; candidate worlds are retained only after rendering and verification against the source video.

4.2 Agentic Discovery Loop

At each round, the agent updates p using evidence η and structured feedback Δ, compiles it into simulator parameters θ, and executes a state trajectory τ containing states, contacts, collisions, and outcomes. For video, verification compares rendered RGB, projected depth, masks, and image-plane tracks at selected frames; for text, it focuses on semantic and physical constraints. The procedure rejects a hypothesis if it does not satisfy the acceptance condition within K iterations.

Shows the shared discovery loop for text and video inputs. Semantic evidence from language or RGB, depth, instance, and track evidence from video constrains an EWR that is proposed, instantiated, executed, rendered, compared, and refined.

Agentic discovery loop for executable world representations from text descriptions or real-world video.
Agentic discovery loop for executable world representations from text descriptions or real-world video.

4.3 Evaluation

For video-driven reconstruction, the authors filter WISA-80K for salient rigid-body motion, excluding clips with substantial camera motion, insufficient object motion, severe editing, or incomplete events, then manually review the remainder. They evaluate up to K = 5 rounds with Visual Alignment, Object IoU, Traj-ADE, Velocity-ADE, and Accuracy@2%D; at an equal five-evaluation budget, the main experiment exceeds Best-of-5 on Visual Alignment, Object IoU, Traj-ADE, and Accuracy@2%D. The appendix reports that, with the physics engine, the fifth round exceeds the matched Best-of-5 baseline on all five metrics.

Demonstrates editability after world construction: changing bowling directions or camera viewpoints changes the simulator rollout and its aligned realistic-video rendering.

Controllable resimulation after changing motion conditions or viewpoints.
Controllable resimulation after changing motion conditions or viewpoints.

Plots reconstruction metrics across one-shot generation and five iterative rounds. At the matched five-evaluation budget, the final iterative result exceeds the Best-of-5 reference for Visual Alignment, Object IoU, Traj-ADE, and Accuracy@2%D, but not Velocity-ADE.

Reconstruction metrics over agentic discovery rounds compared with Best-of-5 sampling.
Reconstruction metrics over agentic discovery rounds compared with Best-of-5 sampling.

4.4 Limitations

The method is limited by simulator coverage: small unobserved differences in terrain, contact geometry, material properties, or other latent conditions can alter rigid-body motion. The loop can therefore converge to an EWR that is locally plausible and visually aligned without providing a mechanistically correct account when the underlying process lies outside the simulator’s modeling scope.

Provides text-to-world examples in which simulation rollouts generated from prompts are transformed into more realistic videos while the paper claims that their encoded trajectories and physical evolution are preserved.

Text-driven construction from executable simulation rollouts to realistic generated video.
Text-driven construction from executable simulation rollouts to realistic generated video.

Provides real-video-to-simulation examples. The abstracted simulations reproduce selected scenes’ coarse composition, motion, and interactions rather than their full photorealistic appearance.

Video-driven abstraction from real observations to executable simulation rollouts.
Video-driven abstraction from real observations to executable simulation rollouts.

5 Application: Learning Quantitative Physical Reasoning

The downstream task predicts a scalar physical quantity from a monocular video and question, covering size, displacement, velocity, and acceleration. World-space questions supply a known reference quantity ρ; following QuantiPhy, the target is calibrated from pixel-space measurement using γ = ρ/ρpix and y = γypix. The paper trains 4B and 9B direct-answer variants with image-space supervision followed by executable-world supervision, and separately trains a 27B reasoning variant.

Illustrates QuantiPhy’s four settings: 2D or depth-aware 3D scenes, each paired with a static or dynamic prior used to calibrate a requested world-space measurement.

Examples of quantitative physical reasoning questions with video evidence and physical priors.
Examples of quantitative physical reasoning questions with video evidence and physical priors.

5.2 Image-Space Measurement Grounding

Image-space supervision is derived from bounding boxes, masks, and tracks in RefCOCO, RefCOCO+, RefCOCOg, RefCLEF, and GOT-10K. Object size and position come from boxes or masks, while motion comes from tracked centers using d = ct2 − ct1, vt = (ct+1 − ct−1)/(2Δt), and at = (ct+1 − 2ct + ct−1)/Δt²; supervised fine-tuning on Dpix establishes localization, measurement, and tracking before metric calibration.

5.3 World-Space Physical Calibration from Verified Executable Worlds

Verified EWRs provide synchronized videos and exact scene geometry, camera settings, timestamps, and physical-state trajectories, enabling world-space VQA labels from Dtext and Dvideo. The second phase uses GRPO with numerical, unit, and format rewards; text-driven worlds supply exact simulator labels, whereas video-driven worlds better match real visual and motion distributions. On QuantiPhy-validation, the 9B direct-answer model reports 55.4 average MRA versus 40.2 for Qwen3-VL-32B-Instruct and 54.8 for Gemini-3.1 Flash; the 27B reasoning model reports 58.6, but its scale and response protocol differ from the direct-answer comparison.

Reports QuantiPhy-validation MRA by 2S, 2D, 3S, and 3D categories. Code-as-World-VL-4B, 9B, and 27B (Reasoning) score 50.6, 55.4, and 58.6 average MRA, respectively; the 27B result is scored from the answer emitted after its reasoning trace.

QuantiPhy-validation quantitative physical reasoning results across proprietary, open-weight, and Code-as-World-VL models.
QuantiPhy-validation quantitative physical reasoning results across proprietary, open-weight, and Code-as-World-VL models.

Places QuantiPhy average MRA against model parameter count. The plotted Code-as-World series rises from 50.6 at 4B to 55.4 at 9B and 58.6 for the 27B reasoning model, while models with undisclosed sizes are grouped separately.

QuantiPhy parameter-scaling comparison using average MRA.
QuantiPhy parameter-scaling comparison using average MRA.

Compares simulator renders with sim-to-real videos. Sim-to-real improves distributional realism from JEDi MMD 3.000 to 1.484 and TRAJAN Fréchet 406.872 to 185.321, while Traj-ADE changes from 1.682 to 1.677 %D, Velocity-ADE worsens from 0.404 to 0.472 %D/step, and Accuracy@2%D falls from 78.81% to 77.49%.

Distributional realism and motion fidelity for text-driven simulator renders and sim-to-real videos.
Distributional realism and motion fidelity for text-driven simulator renders and sim-to-real videos.

Repeats iterative discovery with the physics engine rather than the animation engine. Across five rounds, Visual Alignment, Object IoU, and Accuracy@2%D increase while Traj-ADE and Velocity-ADE decrease; the final result exceeds matched Best-of-5 on all five metrics.

Five-round agentic discovery evaluation with the physics engine.
Five-round agentic discovery evaluation with the physics engine.

Reports pixel-level MRA across four referring-expression datasets and GOT-10K. Adding executable-world reinforcement learning improves both sizes on every listed benchmark; for example, the 9B model rises from 63.7 to 68.3 on RefCOCO and from 20.1 to 26.6 on GOT-10K.

Pixel-level object measurement results for Image-Space and full Code-as-World-VL variants.
Pixel-level object measurement results for Image-Space and full Code-as-World-VL variants.

Qualitatively compares grounding, length, speed, and acceleration predictions. In the selected examples, the Code-as-World-VL outputs are closer to the displayed target boxes and numerical ground truth than the shown comparison outputs.

Qualitative comparison of image-space grounding and measurement predictions.
Qualitative comparison of image-space grounding and measurement predictions.

Separates the contributions of image-space data Dpix and text- and video-driven world data Dtext and Dvideo. Each world-data source improves the listed average over image-space-only training, and using both sources gives the strongest tabulated averages: 50.6 for 4B and 55.4 for 9B.

Ablation of image-space, text-driven-world, and video-driven-world training data.
Ablation of image-space, text-driven-world, and video-driven-world training data.

Displays a video-driven soccer-ball reconstruction with portions of scene.json and simulation_sdk.py. The color coding maps structured fields and SDK operations to composition, evolution, and appearance, and the implementation executes the scene in MuJoCo.

Structured EWR code and simulation interface for a video-driven scene reconstruction.
Structured EWR code and simulation interface for a video-driven scene reconstruction.

Appendix

  • The appendix specifies 73,335 Image-Space question-answer pairs and an executable-world dataset with 1,585 text-driven and 988 video-driven VQA samples. It details MuJoCo-based animation and physics engines, QuantiPhy’s 159 validation pairs, MRA evaluation, reasoning-model settings, realism and motion metrics, and data-source ablations. The ablation table shows that combining Dpix, Dtext, and Dvideo yields the strongest listed QuantiPhy averages: 50.6 for 4B and 55.4 for 9B.

Brief Thoughts

The central contribution is not only simulation-data generation: it makes physical assumptions inspectable, editable, and testable through an executable representation. The reported evidence supports gains on constrained rigid-body reconstruction and QuantiPhy calibration, but broader claims about physical intelligence remain contingent on simulator fidelity, coverage of broader physical regimes, camera-motion robustness, and learning the discovery loop itself rather than only from its verified outputs.