Keynote - The Next Frontier—AI & the Physical World
Topi Manu — AI GTM Lead - Nordics, Google
AI is no longer confined to screens—autonomous systems, AI-powered robotics, and multimodal models are revolutionizing industries from manufacturing to logistics. What does this mean for the future of automation and human-AI interaction?
"Yeah, I've been waiting for this. So if we can talk to dolphins, I think it's a good manifestation that we can cross the boundaries between the physical world and the AI."
Summary
- The talk focuses on the intersection of AI and the physical world through demos and videos. - Gemini Robotics is highlighted for its interactive, dexterous, and general capabilities in real-world tasks. - The speaker demonstrates AI's ability to process real-time data and execute tasks across various modalities. - A new AI model, Dolphin Gemma, is introduced to understand dolphin vocalizations. - The speaker emphasizes Google's comprehensive AI stack and invites collaboration.
Article
AI bridges into the physical world, from robotic dexterity to dolphin communication
Google's Topi Manu unveils next-generation AI systems that cross the digital-physical barrier at VERGE AI Frontier conference
As artificial intelligence systems increasingly interact with our physical environment, the line between digital reasoning and real-world action continues to blur. This emerging frontier took centre stage in Helsinki as Topi Manu delivered a compelling keynote at the VERGE AI Frontier conference, demonstrating how Google's latest AI systems are transforming robotics, real-time computing, and even interspecies communication.
Robots that think and adapt in real time
Manu began his presentation by showcasing Gemini Robotics, Google's advanced vision, language, and action model designed to bring flexible intelligence to physical machines. Unlike traditional robotics that follow predetermined movement patterns, these systems demonstrated an ability to reason about what they perceive and adapt accordingly.
"Many robots can execute predefined actions, but these movements are not predefined. The robot is reasoning both about what it sees and how to move," Manu explained, as video demonstrations showed robots performing complex tasks like origami folding and precise object manipulation.
One particularly striking capability was the system's interactive nature. "To be helpful, robots need to be interactive, responding live to your actions and your voice," Manu noted as he showed robots adapting to changing verbal instructions and environmental conditions in real time.
From code to chess in seconds
The presentation shifted to practical demonstrations of AI in action, with Manu challenging Gemini 2.5 Pro to code a Finland-themed chess game complete with culturally relevant emoji pieces and the Finnish flag's colors. The system produced functional code almost instantly.
"This is the speed with which you can code a new application with the capabilities shown just before," Manu observed, highlighting how AI-assisted development dramatically accelerates the creation of complex applications.
Multimodal intelligence bridges digital and physical worlds
Perhaps the most impressive demonstrations involved real-time visual processing, with the AI system identifying Stockholm from a single image and analyzing physical objects on the fly. This multimodal capability—processing text, images, video, and audio simultaneously—represents a critical advancement for systems that must interpret the physical world.
Manu emphasized that these capabilities extend far beyond chatbots or simple image recognition, enabling AI to reason across different types of information and engage with physical reality in increasingly sophisticated ways.
From human tools to cross-species communication
In what may have been the most ambitious segment of his talk, Manu introduced "Dolphin Gemma," an AI model designed to analyze and potentially decode dolphin vocalizations. This initiative represents an extraordinary extension of AI's capabilities—from enhancing human productivity to potentially bridging communication gaps between species.
"When you stop and think and actually go through what we have, it's way beyond than anybody else and only one with the full stack to be optimized," Manu stated, referencing Google's comprehensive AI infrastructure that makes these diverse applications possible.
The integration challenge ahead
As the session concluded, the scale of both possibility and challenge became clear. The demonstrations showcased AI's growing ability to perceive, reason about, and act within the physical world—a fundamental shift from systems that merely process information to those that can meaningfully engage with our environment.
The journey from robots performing dexterous tasks to potential cross-species communication illustrates the remarkable breadth of this technological frontier. As these systems continue to evolve, they raise profound questions about how AI will transform our relationship with the physical world—questions that go far beyond technical capabilities to the very nature of intelligence and its role in society.
For the audience in Helsinki, Manu's presentation offered not just a glimpse of cutting-edge technology, but a window into a future where the boundaries between digital intelligence and physical reality grow increasingly permeable. The next frontier, it seems, lies not in silicon alone, but in its integration with the world beyond our screens.
Part of VERGE | The AI Frontier: Creativity, Security & Collaboration