When Does Refreshing an HPC System Make More Sense Than Replacing It?

There comes a point in the life of any HPC system when attention inevitably turns to what comes next.

Hardware ages. Research requirements change. Storage fills. Software moves on. New standards emerge. Eventually, parts of the environment may no longer match what is required of them.

It is tempting to think of this as a single lifecycle: the cluster gets old, the cluster gets replaced.  In reality, an HPC system does not age all at once.

The infrastructure supporting a research computing service can change considerably without requiring everything its users depend on to change with it. And sometimes, keeping part of that environment stable is exactly the point.

Stability for users, change underneath

New compute is valuable when researchers need capabilities that their existing environment cannot provide. But when existing CPU and GPU resources still support the workloads being run, changing them just because other parts of the system need attention doesn’t necessarily improve the service.

For researchers, continuity can have value of its own.

A stable compute environment means established workloads continue to run on familiar resources. Existing workflows do not have to change simply because the infrastructure supporting them has reached a different stage in its lifecycle.

Behind that user-facing experience, however, requirements for an HPC service continue to evolve. Security standards change. Operating systems and software need to remain current and supportable. Storage requirements develop. Networking technology moves forward. Power efficiency becomes increasingly important. Core infrastructure has to remain reliable and capable of supporting both today’s service and what may be introduced next.

That creates an important distinction when considering an HPC refresh. The parts of the system that need to change to maintain a modern, reliable service are not necessarily the parts that no longer meet researchers’ needs.

Modernising the service without changing everything

Selective modernisation allows those two requirements to be considered separately.  Rather than assuming the entire platform has reached the end of its life, a refresh can focus on where change will actually improve the service.

That might mean replacing storage or networking infrastructure. It might require an updated operating system and software environment. Core components may need replacing to improve reliability, security, efficiency or supportability.

At the same time, useful compute resources can remain part of the platform where they continue to meet requirements.  This is not about keeping older technology for as long as technically possible. Nor is it about avoiding investment.  It is about directing investment towards the parts of the system where change delivers value, while avoiding unnecessary disruption elsewhere.

What this can look like in practice

City St George’s, University of London, recently took this approach with its Hyperion HPC cluster.  The project retained established CPU and GPU compute resources while modernising the infrastructure surrounding them.

Working together, the City and Alces teams refreshed the surrounding infrastructure, including upgrades to storage, networking, core infrastructure, the operating system and software environment. The work also built on City’s existing ConcertIM managed subscription, providing continuity throughout the refresh as existing compute resources were integrated into the modernised platform.

The resulting changes were significant.

Available usable storage capacity more than doubled. Updates to MPI and filesystem software improved workload performance by up to 10%, while infrastructure changes delivered an approximately 10% improvement in power efficiency.

The refreshed platform also provides a modern software environment that supports current HPC and AI applications while maintaining compatibility with future processor, accelerator, networking, and storage technologies.

The important point is not that City St George’s retained existing hardware. Instead, the team focused investment on the infrastructure around that hardware, updating it so the platform as a whole could continue to provide a reliable, current and extensible research computing service.

Thinking Differently About Infrastructure Lifecycles

Research computing services have always had to balance change with continuity. A refresh creates an opportunity to make that balance deliberate: maintaining stability where it continues to serve researchers while investing where change will improve the service.

Sometimes the right answer will be a completely new platform. New research requirements, capacity demands or technology opportunities can make wholesale replacement the sensible choice. In other cases, modernising the infrastructure around existing compute can provide a more appropriate route forward.

Neither should be the default.

The decision comes from understanding what users need, what the service needs to change, and what the platform will need to support next.

So before planning the next HPC lifecycle around “What should replace our system?”, there may be a more useful place to start:  What actually needs replacing?

Thinking about your own HPC environment?

Whether you’re considering an infrastructure refresh or want to build confidence in delivering a comprehensive, modern HPC user experience through training and enablement, contact us. We can help you explore the options and connect you with expertise from the teams at Alces and ConcertIM.

Get in touch →

Wait, there’s more...

Discover our other blog posts