Google DeepMind and Harvard propose vision-first path to AGI

1 hour ago 23

For years, the race to build artificial general intelligence has been dominated by one modality: language. GPT, Claude, Gemini. All of them learned to reason by digesting oceans of text. A new white paper from researchers at Google DeepMind, Harvard, and other leading institutions argues that approach might be incomplete, and that the path to AGI could run through what machines see rather than what they read.

The paper, titled “Visual General Intelligence: A White Paper” and published on arXiv as 2608.25924, lays out a research agenda for what the authors call visual general intelligence, or VGI. The core thesis: AI systems should learn directly from images, videos, and geometric data to understand, predict, and act in the physical world.

What the paper actually says

This is not a product announcement or a benchmark-beating model reveal. It is a position paper, a collective argument from more than 21 researchers about where the field should invest its attention next.

Among the contributors are Robert Geirhos from Google DeepMind and Yilun Du from Harvard. The work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop, one of the premier gatherings in the computer vision community.

Rather than presenting a single model or architecture, the paper discusses principles, benchmarks, and learning paradigms for building intelligence through visual experience. The researchers argue that generative video models and self-supervised learning, where systems train themselves by predicting what comes next in visual sequences, could provide a foundation for AGI that language alone cannot.

The paper also explores strategies for integrating multiple modalities. Vision and language aren’t positioned as competitors in this framework but as complementary channels, each capturing different slices of intelligence.

Building on DeepMind’s AGI framework

The white paper doesn’t exist in isolation. It builds on a lineage of publications from Google DeepMind that have attempted to map the road to AGI and beyond.

In 2024, DeepMind published “Levels of AGI,” a paper that proposed a taxonomy for measuring progress toward general intelligence. More recently, in June 2026, DeepMind followed up with “From AGI to ASI,” which explored what might come after general intelligence: artificial superintelligence. The VGI white paper slots neatly into this intellectual arc, asking whether the current text-heavy paradigm is sufficient to reach the milestones those earlier papers described.

For now, no specific performance claims, model timelines, or product roadmaps accompany the paper. It is deliberately open-ended, designed to spark exploration rather than declare victory.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article