The University of Oxford granted OpenAI access to historical texts housed in the Bodleian Library for use in training artificial intelligence systems, according to confidential institutional records reviewed by the Guardian. Journalists Ethan Penny and Dan Milmo published their investigation on Saturday, revealing that scanned documents from the collection entered OpenAI's training pipeline.

Oxford announced its partnership with OpenAI in March 2025, framing the collaboration as a means to digitise rare materials and expand access for researchers and students. The university's initial public statements did not mention that the scanned content would serve as training material for AI models.

By June 2025, the Bodleian had transferred 125,000 digitised copies of doctoral theses to OpenAI, according to the Guardian's reporting. These theses originated from academic institutions across Europe and North America during the nineteenth and twentieth centuries.

Internal meeting notes obtained through freedom of information procedures indicate that some Bodleian staff members harboured reservations about the arrangement. Concerns centred on potential damage to the institution's reputation and the environmental footprint associated with AI model training.

Oxford's administration countered these objections by emphasising that the scanned materials are in the public domain, represent a modest portion of the collection and are not exclusively licensed to OpenAI. A university representative stated that the library retains ownership of the scans and plans to make them publicly available online within the coming months.

The spokesperson further asserted that the training application had not been concealed from staff. While digitisation constituted the primary objective, university personnel had been informed from the outset that the texts would also contribute to model development.

With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives

OpenAI spokesperson, to the Guardian

Oxford holds the distinction of being the sole British institution participating in OpenAI's NextGenAI initiative. The consortium also includes Boston Public Library, Caltech, MIT and the University of Michigan.

The arrangement reflects a broader trend among AI developers seeking high-quality training datasets from physical books, as the internet increasingly contains machine-generated content. Some companies have resorted to destructive scanning methods, dismantling volumes to extract text, a practice that has generated friction within the used book trade.

In August, 404 Media documented the routing of rare books through an Amazon facility designed to scan and discard volumes for AI purposes. The Guardian noted that Oxford's agreement preserves the physical integrity of the Bodleian's collection.

Source: The Next Web