University of Oxford allows OpenAI to train AI models on Bodleian library data

1 hour ago 27

One of the world’s oldest universities just handed one of the world’s newest tech giants the keys to its library. The University of Oxford has entered a five-year partnership with OpenAI, allowing the company behind ChatGPT to train its AI models on historical texts housed in the legendary Bodleian Libraries.

The collaboration, announced on March 4, 2025, kicked off with a pilot project in February targeting the digitization of roughly 3,500 public-domain dissertations dating from 1498 to 1884. These are texts that have sat on shelves for centuries, largely inaccessible to anyone who couldn’t physically visit Oxford.

What the deal actually involves

The partnership falls under OpenAI’s NextGenAI program, which has committed $50 million in research grants, computing resources, and API access to a handful of elite institutions. Oxford joins Harvard and MIT in that cohort.

For the Bodleian Libraries, the pitch is straightforward: AI-powered metadata enrichment and improved transcription services will make centuries-old collections searchable and globally accessible.

On the education side, Oxford plans to roll out ChatGPT Edu across the university starting in September 2025. The tool has already been piloted with about 750 users. Research findings from the digitization project are expected in open-access reports by early 2026.

Why staff aren’t thrilled

Not everyone at Oxford is celebrating the arrangement. University staff have voiced concerns about the reputational risk of partnering with OpenAI, a company that has faced sustained criticism over copyright issues, data sourcing practices, and its rapid commercialization under CEO Sam Altman.

The dissertations in question are public domain, meaning no copyright applies. But the broader optics of a 900-year-old academic institution feeding its collections to a for-profit AI company have understandably raised eyebrows.

There’s also the question of what OpenAI gets versus what Oxford gets. The university receives digitization support, compute resources, and educational tools. OpenAI receives high-quality, curated training data from a globally trusted source, plus the implicit endorsement that comes with an Oxford partnership.

The bigger picture: AI’s hunger for academic data

Oxford’s deal with OpenAI fits a pattern that has been accelerating across higher education. Tech companies are increasingly turning to academic institutions as data sources, partly because web-scraped data is running into legal challenges and partly because academic corpora tend to be higher quality than the average internet text.

The NextGenAI program’s $50 million commitment, spread across multiple institutions, is substantial but modest relative to OpenAI’s overall spending, which has been measured in billions annually.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article