Oxford Lets OpenAI Train on 125,000 Bodleian Scans as AI Firms Hunt Fresh Data
Updated
Updated · The Guardian · Sep 26
Oxford Lets OpenAI Train on 125,000 Bodleian Scans as AI Firms Hunt Fresh Data
3 articles · Updated · The Guardian · Sep 26
Summary
Internal Oxford documents show OpenAI used digitized Bodleian material to “populate” its training set, a use not stated in the university’s March 2025 partnership announcement.
125,000 images from historical dissertations had been shared by June 2025, and scans also include 10,000 16th-century broadside ballads, with staff discussing more archives such as Irish state papers and Dorothy Hodgkin notebooks.
Oxford said the project is modest, limited to out-of-copyright works, non-exclusive for OpenAI, and will put the scans online within months while the Bodleian keeps rights and its 23 million-item collection remains intact.
Meeting minutes show staff raised reputational and environmental concerns, even as OpenAI and other developers increasingly turn to libraries and physical books because web data is being polluted by AI-generated content.