← The Vault
Explainer

Why AI companies are desperate to give robots common sense

While tools like ChatGPT are experts at manipulating words, they struggle to grasp how the physical world actually works. New startups are now racing to build 'world models' that help robots predict what happens when you nudge a glass or walk down a street. This shift from pure text processing to physical awareness is the next major frontier for artificial intelligence, turning the focus away from abstract intelligence toward real-world competence.

Edition № 237Room: Explainer16 July 20262 min readSources: 6
Article

Digital assistants like ChatGPT are essentially advanced autocomplete engines. They are highly efficient at predicting the next word in a sentence, but they exist entirely within the realm of text and code. They have never felt the weight of an object or navigated a crowded room. Because of this, artificial intelligence has hit a wall when it comes to the physical world, leading industry researchers to focus on a new goal: building machines that understand how reality functions rather than just how language works.

WHAT'S HAPPENING

Experts are moving away from just trying to make AI smarter at reading and writing. Instead, they are pouring massive investments into what are called world models. A world model is an artificial intelligence designed to predict how the physical environment will change over time. If you were to push a glass off a table, you intuitively know it will fall and likely shatter. A world model aims to give a computer that same kind of physical intuition. This is fundamentally different from the current generation of tools, which rely on massive databases of text to simulate reasoning. Today, companies ranging from well-funded startups to established labs are racing to bridge this gap, often partnering with robotics and manufacturing firms to train their AI in factories and workshops where physical interaction is constant.

Giving AI a sense of reality

HOW IT WORKS

To understand why this is a new challenge, think of a doctor who has memorized every medical textbook ever written but has never actually stepped inside a hospital. That doctor could likely answer any question you ask about biology, but they would be helpless if they had to perform surgery or adjust to an emergency in a room full of moving people. Current language-based AI works the same way: it is trained only on text, which is an abstract representation of reality. A world model, by contrast, is designed to observe and predict the state of the world itself. It learns by processing data from sensors—like cameras or robotic touch pads—so it can predict what happens next when an object moves or an environment shifts. While a language model predicts your next word, a world model predicts the next state of the physical scene in front of it.

WHY IT MATTERS

This shift is essential if we ever want to move beyond screen-based assistants and into a world where machines can function safely around us. A robot in a house or a public park cannot rely on textbooks to interpret its surroundings; it needs to understand the consequences of its movements. By prioritizing physical awareness, researchers are acknowledging that intelligence isn't just about processing information—it is about acting reliably in an unpredictable environment. Whether this leads to smarter factory automation or truly capable household robots, the goal is the same: software that doesn't just think, but understands the consequences of existing in the real world.

Sources
← PreviousWhy companies worry about their AI despite using it every dayNext →When AIs attack: A new kind of cybersecurity incident
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault