22 September 2026 - From Storage Shock to AI-Ready: What Imperial College learned
By Adrian Mannell, Digital Workplace Technology Lead, Imperial College London
When Microsoft changed its M365 storage model, Imperial College London faced a number that stopped us in our tracks: a projected 3.6 petabytes of data on a trajectory that was no longer sustainable. With 629 million OneDrive files in scope, this was not a tidy housekeeping exercise. It was a governance crisis hiding in plain sight - and, as it turned out, an unexpected opportunity.
This is not a story about a perfect transformation. This is what happens when a research-intensive university is forced to ask a question it was quietly avoiding: Does our current platform support research data and address the challenges around growth, governance and ownership?
The real problem was never the quota
The quota change was the trigger, not the cause. What it exposed was something more fundamental: personal storage had become long-term research storage. We had set no clear boundaries between personal work and institutional research storage, so users naturally saw OneDrive as the default place for both. Ownership was unclear. Retention policies were inconsistent. And the question had shifted, almost imperceptibly, from 'how do we make space?' to something far more consequential: how do we properly govern and sustain our data for the long term?
A single research programme at Imperial might need fast compute storage, external collaboration, secure processing for sensitive clinical data, and long-term preservation for citable outputs - all at the same time. No single platform serves every stage equally. Yet for years, the path of least resistance had been OneDrive, regardless of whether it was the right fit.
The lesson we took from this: treat the quota event as a governance moment, not only a clean-up exercise.
Acting in Four Lanes Simultaneously
Policy alone was too slow. Migration alone would simply move the problem. We needed to act across four lanes at once:
- Discover (tenant-wide analysis, large-account targeting, version and activity insights)
- Constrain (phased quotas, headroom limits, clear accountability)
- Support (segmented communications, self-service guidance, assisted migration)
- Move (user-led, bulk transfer, and pre-staged shadow migrations).
The difference was orchestration - technology, governance, communication, and support working in parallel, not in sequence.
We designed three distinct migration journeys rather than one process for everyone. Self-service for confident users with manageable volumes. Assisted bulk transfer via Box Shuttle for large or complex accounts.And pre-staged shadow migration for the hardest cases - copy first, cut over when required. We needed an environment that would actively support our researchers and future-proof exponential data growth without penalising them with sudden caps. Box provided an unlimited storage solution to meet that research demand, enabling us to target approximately 1.5 petabytes moved across. We were candid with ourselves: early take-up for assisted migration was low. Hard quotas and targeted prompts eventually changed behaviour. The operational lesson is to design for the user, not the technology.
Making controls travel with the content
The governance work that followed was, in many ways, more important than the migration itself. We aligned Imperial's Box classification labels to our Microsoft Purview set, creating a consistent vocabulary across platforms. The principle: governance is the bridge between "stored" and "safe to use with AI."
Our governance fabric operates across four dimensions.
- Classify - one university vocabulary across Purview and Box.
- Protect - sharing controls, access management, threat detection, and secure spaces.
- Retain - lifecycle rules aligned to research and records needs.
- Assure - audit, reporting, ownership, and exception handling.
Classification maturity takes time. We were honest about that. Threat detection can create false positives, and exceptions need a governed route. Waiting for a perfect classification standard would have missed the quota deadline. Success came from creating enough control to act safely while the longer-term model matured.
AI readiness is a content quality outcome
Here is the reframe that changed how we think about this entire programme: AI amplifies the quality, permissions, and context of the content it can reach. That means AI readiness is not a technology project. It is a content quality outcome.
The journey runs: consolidated and findable content → governed (classified, retained, access-controlled) → trusted (high quality, current, auditable) → AI-ready (eligible, grounded, purpose-limited). Do not connect AI to everything simply because you can. The goal is not maximum retrieval. It is trustworthy retrieval within purpose, permission, and governance boundaries.
This creates a concrete benefit from the storage and classification work that goes well beyond cost avoidance.
dAIsy: Equity of Access to AI
One of the most important decisions Imperial made was to address the equity problem in AI access. Department-by-department purchasing creates haves and have-nots. Our answer is dAIsy - a universal AI access tool built on NebulaOne - providing broad access to multiple approved AI models without a two-tier licence culture.
dAIsy is universal (available across the university), multi-model (choice of approved model capabilities), controlled (central access, token and spending controls), protected (content handling within approved boundaries), and evolving (a platform for agents and new models over time). The principle: safe baseline access for everyone, specialist tools where value and risk justify them.
What we would tell our peers
- Right now: We are finishing the quota transition, moving data across, and guiding users into dedicated research spaces.
- Next up: We are automating governance, letting classification labels automatically manage sharing permissions, access controls, and retention.
- As we build momentum: With clean, governed data foundations established, we are connecting trusted content repositories to pilot AI capabilities and approved retrieval models safely.
- At scale: We will expand based on proven research outcomes, validated risk controls, and sustainable value.
The UK higher education takeaway is this: storage strategy, content governance, and AI strategy are now one agenda. The institutions that treat them separately will find themselves rebuilding the foundation when AI demands it. The institutions that treat them as one will find they have already done the work.
Three asks to peers:
- Establish dedicated research spaces with clear placement guidelines
- Make classification operational
- Provide a safe default AI route.
Do not wait for perfect metadata. Start with high-value, lower-risk teams and lines of business, and measure outcomes as you iterate.
Want to hear more? Hear from Adrian Mannell, Digital Workplace Technology Lead at Imperial College London and Anthony Prior, EMEA Higher Education Lead, Box as they discuss more about Imperial’s story HERE. Want to meet Box in person? You’ll find them in person over on Stand 9 at DISC26 on 7-8 October and Box Summit London on December 3.
