Code-as-World Could Change How SMEs Use Video AI

Code-as-World Could Change How SMEs Use Video AI — featured image

by

Why This AI Story Matters to Your Business

A new approach called “Code-as-World” is turning real video into editable, executable physics programs. Instead of treating a video as a sequence of images, the system attempts to identify objects, movement, dimensions, contact, gravity and other physical relationships, then recreate them inside a MuJoCo simulation.

For you as a Malaysian SME owner, this matters because many business processes already depend on video: checking warehouse activity, reviewing production lines, analysing delivery handling, inspecting machinery and training new staff. Conventional video AI can tell you what appears in a frame. A physics-aware system aims to help explain what happened, what moved, what collided and whether a different action may have produced a better result.

The technology is still at the research and internal-prototype stage, not a ready-made solution for every company. However, it points towards a practical direction: business video may eventually become a source of testable operational models instead of merely an archive for people to watch later.

What Happened

MirroS released Code-as-World, a method that represents a physical scene as executable code rather than only pixels, captions or hidden model features. Its representation combines composition, evolution and appearance. Composition includes objects, geometry, dimensions, mass, friction and gravity. Evolution describes initial states, forces, contacts, collisions and duration. Appearance covers the camera, lighting, materials and rendering settings. The technical report explains that separating these elements allows the physical rules to be changed without changing the visual presentation. Source

The released implementation converts this representation into a scene.json file that can run in MuJoCo. It includes an animation engine for kinematic poses and a physics engine for forces and contacts. This means a generated scene can be edited, executed and rendered again, giving an agent a way to compare its reconstruction with the original footage. Source

Rather than making one prediction, the system uses an agentic loop: propose, instantiate, execute, render and verify. It can repeat this process for up to five rounds. Supporting tools provide object masks, image-plane tracks, depth and camera geometry, as well as per-object meshes. The resulting simulation is compared with selected video frames using colour, depth, masks and trajectories. Differences are converted into feedback that guides the next revision. Source

MirroS reported that the five-evaluation process outperformed five independent samples on visual alignment, object intersection-over-union, trajectory error and accuracy measures. The result also held when tested with a second execution engine. Source

Why This Matters for Malaysian SMEs

Consider a small food manufacturer in Selangor or Penang. A production video may show workers moving trays, boxes or ingredients around a line. A normal computer-vision system could count items or flag that a person entered a restricted zone. A physics-aware reconstruction could eventually help identify why trays toppled, whether spacing was too narrow, or how a change in movement affected the next workstation.

For a logistics SME operating around Port Klang, Johor Bahru or East Malaysia, recorded loading footage could be used to study how parcels shift inside a vehicle. The useful question is not only “which box moved?” but also “what contact or acceleration caused the movement?” If the scene can be recreated in an editable simulator, you could test different stacking arrangements virtually before changing the loading procedure.

Retailers, repair workshops and construction subcontractors may also benefit. A shop could review how customers interact with displays. A workshop could recreate how a component was handled before damage appeared. A site supervisor could analyse the order in which materials were moved. These uses would require careful validation, but the central idea is valuable: convert video evidence into a model that can be inspected and tested.

Reported capability Possible SME relevance
Executable scene.json Create an editable representation of a recorded process
Physics and animation engines Compare actual movement with planned or simulated movement
Up to five propose-and-verify rounds Improve reconstruction through repeated checking
Object, depth and trajectory comparison Review handling, movement and spatial relationships
Open-source 4B and 9B checkpoints Allow technical teams to evaluate the research internally

The most realistic starting point for your company is not a fully autonomous factory. Begin with one repeatable process and one measurable question. For example, you might ask whether a particular packing sequence causes more product movement, or whether a trolley route creates unnecessary contact with shelving. Record the process consistently, define what success means and keep a human reviewer involved.

“Pixels are evidence of a physical scene, not its ontology.” — Core argument reported in the Code-as-World technical work.

This distinction is important. A video can show that a carton fell, but it may not directly reveal its mass, friction, contact force or the exact reason it lost balance. A simulation that claims to explain the event must therefore be checked against the footage and against real operational knowledge.

The Bigger Picture

Code-as-World reflects a broader shift from AI that recognises patterns to AI that builds testable representations. In a business setting, this could support digital twins, safety reviews, staff training and process improvement. Instead of asking an AI assistant to summarise a video, you could eventually ask it to recreate a workflow, identify possible failure points and show how a proposed change affects the sequence.

The research used verified executable worlds as training supervision. Its training process included 73,335 image-space question-and-answer pairs and executable worlds created from 1,585 text-driven and 988 video-driven scenes. Source The reported Code-as-World-VL-9B model achieved 55.4 MRA on QuantiPhy validation, compared with 54.8 for Gemini-3.1 Flash and 40.2 for Qwen3-VL-32B-Instruct, according to the source article. Source

Those results should be read carefully. The system is described as rigid-body focused, and the model does not learn the discovery loop itself. The benchmarks also do not automatically prove that it will understand Malaysian factories, local warehouse layouts or your specific machinery. A strong benchmark result is a reason to investigate, not a reason to automate a safety-critical decision immediately.

For now, Malaysian SMEs can prepare by improving video data practices. Use fixed camera positions where possible, record timestamps, label important objects, document process steps and store incident examples. Ensure employees know when recording takes place and protect footage containing faces, customer information or confidential operations. Good data governance will matter just as much as the AI model.

Code-as-World is best understood as an early signal. The future of business video AI may involve editable simulations that let you test operational changes before applying them in the real world. If you identify a narrow workflow, gather consistent evidence and verify every recommendation with experienced staff, your company can explore this direction without treating an experimental system as an unquestionable authority.

Ready to Streamline Your Operations?

Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →