Is Artificial Intelligence Evolving Into a “Physical-Universe AI”?
From language models to world models, physics foundation models, and agents that can understand, predict, design—and eventually act upon—the physical world.
Artificial intelligence may be approaching another architectural transition. The first great wave of modern foundation models taught machines to manipulate language. Multimodal models expanded that capability to images, sound and video. Agentic systems are adding planning, tools, memory and action. Now a growing group of researchers and companies is pursuing something different: AI that develops an internal representation of space, time, objects, forces, materials, motion, causality and physical processes sufficiently powerful to predict what will happen, simulate alternatives, optimize designs and ultimately control machines operating in the real world.
There is not yet a universally accepted name for this emerging category. Researchers variously call its components physical AI, embodied AI, world models, spatial intelligence, neural-operator models, physics foundation models, robot foundation models, vision-language-action models, and world-action models. The distinction matters because these technologies attack different pieces of the problem. But collectively they suggest a larger possibility: AI may be evolving from models primarily representing human information into systems capable of modeling—and eventually acting intelligently within—the physical universe itself.
That does not mean anyone has created a model of “the universe” in the literal scientific sense. Today’s systems remain specialized, incomplete and experimentally constrained. Nevertheless, the direction of travel is becoming increasingly clear. A recent industry overview notes accelerating investment in world models, while researchers are increasingly investigating systems that maintain representations of space, time and changing physical states rather than relying primarily on language.
The August 2026 unveiling of Accelerated Understanding makes the question particularly timely. Reuters reports that its founders are pursuing an AI architecture based not on the Transformer underlying most modern language models, but on neural operators designed to learn physical phenomena across space and time. The company says potential applications range from semiconductor engineering and robotics to extreme-weather prediction and energy exploration.
It is therefore useful to ask whether these apparently separate efforts represent the beginnings of a broader technological category.
The Emerging Physical-Universe AI Ecosystem
1. — Anima Anandkumar and Benedikt Jenik — California. Founded by Caltech professor and former NVIDIA AI research director Anima Anandkumar and AI-infrastructure engineer Benedikt Jenik, Accelerated Understanding is developing perhaps the clearest example of what could properly be called a physics foundation model. Its neural-operator architecture accepts states of physical systems and attempts to predict their evolution directly through three spatial dimensions plus time, rather than reducing the problem to text or sequences of two-dimensional images. The company reports models as large as one trillion parameters, scaling experiments to 35 trillion parameters, training workloads involving petabytes of physical data and inference exceeding five trillion context elements, although these extraordinary figures remain company-reported rather than independently validated. It also reports that one model can learn multiple areas of physics and exhibit what it calls “cross-physics uplift,” potentially allowing knowledge learned in one physical domain to improve another. Most significantly, it envisions a closed engineering loop of define → simulate → improve → simulate again, potentially allowing AI not merely to predict physical outcomes but to recommend directions for improving a design.
2. World Labs — Fei-Fei Li, Justin Johnson, Christoph Lassner and Ben Mildenhall — San Francisco Bay Area. World Labs approaches the problem through spatial intelligence, with Fei-Fei Li and her co-founders developing foundation models intended to perceive, generate, reason about and interact with three-dimensional worlds. Its first product, Marble, generates persistent, spatially coherent 3D environments from images, videos or text, providing a very different route toward world understanding from Accelerated Understanding’s mathematical-physics approach. The central hypothesis is that genuinely intelligent machines require an understanding of three-dimensional geometry, objects, relationships and environments rather than treating visual experience merely as collections of pixels. Such models could eventually support robotics, simulation, architecture, design, games, digital twins and machines navigating real environments. World Labs therefore represents the spatial-representation branch of physical-universe AI.
3. AMI Labs — Advanced Machine Intelligence — Yann LeCun, Alex LeBrun and an international research team — Paris, New York, Montreal and Singapore. Founded after Turing Award winner Yann LeCun left Meta, AMI is pursuing action-conditioned world models that allow an intelligent system to predict the consequences of possible actions before taking them. AMI argues that intelligence begins in the world rather than in language and intends to combine world models with persistent memory, reasoning, planning and controllability. The company announced $1.03 billion in initial financing in March 2026 and identifies industrial process control, automation, robotics, wearable systems and healthcare among its potential applications. An action-conditioned model could effectively ask, “If I do A, what happens next; if I instead do B, what happens?” That makes AMI especially significant because it connects world modeling directly to agent planning and decision-making.
4. Physical Intelligence — Sergey Levine, Karol Hausman, Chelsea Finn, Brian Ichter and Lachy Groom — San Francisco. Physical Intelligence is developing general-purpose robot foundation models, including its π family, intended to provide robots with capabilities analogous to the generality language foundation models brought to text. Its original π ₀ combined vision, language and physical actions, producing continuous motor commands rather than merely language tokens and operating across multiple robot types. The company’s April 2026 π ₀.7 research demonstrated robots attempting tasks they had not explicitly encountered during training, suggesting early forms of generalization rather than conventional task-specific robot programming. The objective is eventually straightforward but extraordinarily difficult: allow a person to tell a robot what needs to be done and have the machine determine how to accomplish it. Physical Intelligence therefore represents the embodied-action layer of this emerging architecture.
5. Skild AI — Deepak Pathak, Abhinav Gupta and team — Pittsburgh. Skild is developing the Skild Brain, which it describes as an “omni-bodied” general-purpose robotic intelligence intended eventually to control many different robot morphologies rather than being designed around one machine. Its hierarchical architecture separates higher-level manipulation and navigation decisions from a high-frequency control policy that translates those decisions into joint positions and motor torques. Training incorporates simulation, human video, teleoperation and experience generated by deployed robots, creating the possibility of a continuing physical-data feedback loop. The company says its systems are already being applied to security and inspection, mobile manipulation, packing, construction, warehouses, factories and other physical applications. Its $1.4 billion Series C announced in early 2026 demonstrates the scale of investment now entering general-purpose physical intelligence.
6. — Ali Agha and a team drawn from organizations including NASA JPL, DeepMind, Tesla and NVIDIA — Irvine, California. FieldAI is building Field Foundation Models and EDGE, a general-purpose robot-brain platform intended to operate across different robots, tasks and environments. A particularly important element is its Belief World Model, which attempts to reason probabilistically about uncertain environments rather than assuming the robot possesses complete and perfectly reliable information about its surroundings. That distinction becomes critical when AI leaves controlled digital environments and encounters changing terrain, unexpected objects, sensor errors, people and equipment. FieldAI reports deployments across hundreds of complex industrial environments and has raised more than $400 million. Its approach represents an important uncertainty-aware and risk-aware branch of physical AI, especially relevant to industrial, infrastructure and safety-critical applications.
7. — Jensen Huang and NVIDIA’s physical-AI organization — Santa Clara, California. NVIDIA is constructing much of the enabling infrastructure beneath the emerging physical-AI ecosystem rather than pursuing only one end-user robot. Cosmos 3 is described as an open physical-AI foundation model combining vision-language reasoning with world and action generation, while NVIDIA’s broader platform includes Isaac simulation and robotics technologies and GR00T robot models. Cosmos can be post-trained using embodiment-specific data to create World Action Models that connect understanding of an environment with actions a machine could take inside it. NVIDIA is consequently attempting to provide a development stack spanning accelerated computing, simulation, synthetic data, world modeling, reasoning, robot policies and deployment hardware. Its importance is comparable to its earlier role in generative AI: NVIDIA may become an enabling computational platform upon which many independent physical-AI companies build.
8. — Google DeepMind robotics and Gemini research teams — London and Mountain View. Google DeepMind is connecting its Gemini multimodal foundation models with physical action through Gemini Robotics, whose 1.5 generation is a vision-language-action model capable of converting visual observations and human instructions into robot motor commands. The system is designed to reason about physical spaces, break objectives into intermediate steps, adapt behavior when conditions change and accept natural-language corrections while operating. DeepMind’s broader world-model research also includes the Genie family of interactive environment models, making the organization significant on both the simulation/world-generation and embodied-action sides of the field. Gemini Robotics illustrates how existing multimodal intelligence can be extended downward from reasoning into physical control. It is therefore an important example of the evolutionary path multimodal foundation model → agent → embodied agent.
9. Project Prometheus — Jeff Bezos and Vik Bajaj — United States. Prometheus remains unusually secretive, but Reuters reports that Bezos and biotech entrepreneur Vik Bajaj are building an extraordinarily well-capitalized AI venture aimed at automating the manufacturing of complex physical systems. The company reportedly raised a $12 billion Series B in June 2026, after earlier attempting to recruit Accelerated Understanding founders Anandkumar and Jenik into senior scientific roles. Public technical details remain insufficient to determine whether Prometheus is developing one physics foundation model, multiple specialized scientific models, autonomous laboratories, manufacturing agents or some combination of these technologies. Nevertheless, its stated emphasis places it at the intersection of AI, engineering, experimentation and manufacturing rather than conventional language-oriented generative AI. Its financing alone makes Prometheus one of the most consequential organizations to watch as AI moves from manipulating information toward creating and manufacturing physical things.
Are These Really One New Category?
[Likely] Yes—but the category is still forming.
It would be premature to describe all these technologies as a single architecture. Accelerated Understanding’s neural operators, World Labs’ spatial models, AMI’s action-conditioned world models, NVIDIA Cosmos, robot foundation models and vision-language-action systems solve substantially different problems.
Yet their convergence is striking.
One possible future stack now looks like this:
Language and multimodal intelligence
→ understands instructions, knowledge and objectives
Spatial intelligence
→ understands objects, geometry and three-dimensional environments
Physics foundation model
→ predicts forces, fields, materials and physical evolution
World model
→ predicts how an environment changes
Action-conditioned world model
→ predicts what happens if the agent performs a particular action
Planning and reasoning agent
→ evaluates alternative courses of action
Engineering/scientific agent
→ simulates, experiments and optimizes
Embodied foundation model
→ translates decisions into physical actions
Robotic or manufacturing system
→ changes the physical world.
That is substantially more than today’s familiar chatbot.
From Artificial Intelligence to Artificial Physical Intelligence
If this convergence succeeds, its importance will come from closing a fundamental gap in contemporary AI. Language models have accumulated extraordinary representations of what humanity has written about the world. Physical-universe AI would increasingly represent how the world itself behaves.
An engineer could eventually give an AI agent a performance requirement rather than a completed design. The system could propose an architecture, simulate thermal, structural, fluid and electromagnetic behavior, identify weaknesses, modify the design, simulate it again and present engineers with optimized alternatives. Accelerated Understanding explicitly describes the beginnings of this simulate-and-improve loop.
Manufacturing could extend the process further:
requirement → design → physics simulation → optimization → manufacturing planning → robotic production → inspection → operational feedback → redesign.
Science could undergo a comparable transformation. Instead of AI merely searching papers or suggesting hypotheses, scientific agents could combine accumulated scientific knowledge with physical models, propose experiments, simulate candidate experiments, determine which uncertainties actually require physical testing, interpret measurements and revise the underlying hypothesis.
Energy systems could model reservoirs, turbines, electrical networks, weather and materials. Semiconductor engineers could explore enormous combinations of geometries, materials, thermal characteristics and electromagnetic effects. Aerospace engineers could evaluate aerodynamics, propulsion, structures and controls. Materials scientists could search candidate structures before synthesizing them. Robotics could give machines an ability to anticipate physical consequences rather than simply react to sensor inputs.
Security introduces an equally consequential dimension. A sufficiently capable physical agent might reason about buildings, transportation networks, industrial processes, electrical systems, communications infrastructure and autonomous machines. The same capabilities that make these systems powerful therefore make identity, authority boundaries, simulation validation, consequence prediction, human approval, runtime monitoring and containment increasingly important. Physical AI converts an incorrect inference from a potentially bad answer into a potentially bad physical action.
That suggests an important distinction for the next stage of AI development:
Intelligence tells the system what might be true.
Reasoning tells it what might be done.
World models tell it what might happen.
Physics models tell it why and how the physical system will change.
Agents decide what actions to pursue.
Physical AI gives those decisions consequences in the real world.
[Likely] We therefore may indeed be witnessing the beginnings of a broader Physical-Universe AI capability—not a single new model replacing the LLM, but an additional set of foundation-model and agent capabilities being assembled around it.
The long-term destination would be considerably more important than simply producing better robots. It would be AI capable of helping human beings understand, predict, simulate, design, optimize, manufacture and safely operate increasingly complex portions of the physical world.
That would represent a fundamental expansion in what artificial intelligence is for: from helping humanity process its accumulated information to helping humanity understand and deliberately reshape the physical systems upon which civilization depends.
- Log in to post comments