Some systems are experiencing issues

About This Site

This site will show any outages being experienced by the ilifu system.

Documentation for using ilifu is available here: https://docs.ilifu.ac.za/

Please log any issues you may experience using our support email address support@ilifu.ac.za

Stickied Incidents

1st September 2026

CephFS mounted Filesystems ilifu facility offline for CephFS metadata recovery starting 2 Sept, 15:00

We have made the difficult decision to take the ilifu facility offline in order to complete a more extensive CephFS metadata recovery process. To resolve the current issues, several scans of the CephFS data pool must be conducted while CephFS is offline. Due to the large number of objects the scans themselves are expected to take several days.

We will shutdown the ilifu services at 15:00 SAST tomorrow, 2 September, and we expect ilifu to be offline for several days. Please copy any critical scripts or files you may need to your local machine before the shutdown tomorrow afternoon. You will not have access to your data while the system is offline.

  • The ilifu services and CephFS are again operational.

  • We are pleased to confirm that ilifu services will return online at 12:00 SAST on Wednesday, 16 September 2026, including access to the production CephFS filesystem.

    If you are using the temporary NFS service, please finish your running jobs and log out before 12:00 on Wednesday. Your data on this temporary storage will remain available for a few weeks so that you can transfer it to your usual working area on CephFS. We will communicate the migration deadline and arrangements separately.

    Following extensive recovery and verification work, a small number of metadata inconsistencies remain, affecting a very small fraction of the approximately 1.1 billion files on ilifu. Further recovery attempts for these remaining cases would require keeping ilifu completely offline and unavailable for all users for another 3-4 weeks, with a low probability of success. Having assessed the likely benefit against the impact of extending the outage for all users, we have decided to restore service and address the few outstanding issues separately. For those small numbers of users impacted, we hope you can understand why we have come to this decision. We will be reaching out to those we have identified as being impacted separately, however, if you encounter any file system anomalies, please report them to support@ilifu.ac.za.

    This outage resulted from a highly unusual and unfortunately timed confluence of factors. It has reinforced the need for stronger operational controls, responsible use of shared resources by users and project leads, and additional redundancy measures. We will be reviewing our file-transfer policy and the approval process for research groups, particularly for transfers involving large numbers of files. Storage allocations will be enforced more strictly: once access is restored, please review your usage against your project allocations and delete what you don't need - you'll be improving ilifu's overall resilience in the process.

    We again apologise for the extended outage and recognise the disruption and stress it has caused researchers, particularly students and postdocs approaching thesis or paper submission deadlines. We appreciate your patience and the sustained efforts of the recovery team. Working together, we can strengthen the resilience of ilifu while supporting the diverse needs of our research community. We also look forward to sharing news of new projects and expanded support in the coming months.

  • We are still working on the cephfs metadata recovery, and cephfs existing filesytems have been taken offline. We have bought up new (empty) /users folders in the interim so that people are able to login and work. NB: Old data will not be accessible at the moment while repairs are conducted.

  • 21st August 2026

    Slurm GPU Nodes GPU nodes allocated to outreach event

    The GPU nodes, gpu-006 and gpu-007, will be unavailable in ilifu Slurm from 17 September - 11 October (previously incorrectly indicated 5 October), as these resources will be dedicated to an outreach event supported by IDIA.

    5th June 2026

    Access to ilifu cloud infrastructure

    Dear colleagues

    As hosts of the ilifu cloud infrastructure, the University of Cape Town (UCT), in coordination with the Inter-University Institute for Data Intensive Astronomy (IDIA), are issuing this advisory because SSH keys used to access ilifu may have been compromised. If you use SSH keys for access, please rotate them immediately and replace the old public keys on all systems and services where they are trusted.

    SSH keys are commonly used across servers, code repositories, cloud platforms and automation systems. If a private key is exposed, an unauthorised person may be able to authenticate as you on any system that trusts the matching public key.

    Please note that SSH private keys are not limited to the computer where they were created. If a private key is compromised, it may be used from another system wherever the corresponding public key is still trusted.

    Required action

    If you use SSH keys, please complete the following steps:

    1. Generate a new SSH key pair.
    2. Add the new public key to every system or service where you use SSH authentication.
    3. Test that the new key works as expected.
    4. Remove the old public key from all systems and services where it was previously trusted.

    Systems and services that may be affected include Linux or Unix servers, bastion or jump hosts, code repositories such as GitHub, GitLab or Bitbucket, cloud platforms, automation or deployment systems, backup or monitoring systems, and any system where your public key appears in an authorised_keys file.

    Best practice guidance

    Please follow these SSH key hygiene practices:

    • Use separate SSH keys for different purposes, such as work, personal use, production access and code repositories.
    • Protect private keys with a strong passphrase.
    • Never share private keys with anyone.
    • Never send private keys by email, chat, ticketing systems or file-sharing platforms.
    • Do not store private keys in scripts, shared folders, source code repositories or documentation.
    • Remove old or unused public keys from systems and services.
    • Avoid SSH agent forwarding unless specifically required and approved.
    • Review your ~/.ssh/config, shell history, Git remotes, scripts and automation files for references to old keys.
    • Ensure private key files have restrictive permissions and are readable only by your user account.

    What to do with old keys

    After your new key is working, remove the old public key from every system or service where it was configured. Deleting the old private key from your workstation is not enough; the old public key must also be removed everywhere it was trusted.

    If you are unsure where your SSH key is used or need assistance rotating it, please contact the relevant support team as soon as possible.

    Should you have any questions relating to this email, please direct your query to idia-security@uct.ac.za.

    Regards

    University of Cape Town, in coordination with the Inter-University Institute for Data Intensive Astronomy

    Past Incidents

    1st September 2026

    CephFS mounted Filesystems ilifu facility offline for CephFS metadata recovery starting 2 Sept, 15:00

    We have made the difficult decision to take the ilifu facility offline in order to complete a more extensive CephFS metadata recovery process. To resolve the current issues, several scans of the CephFS data pool must be conducted while CephFS is offline. Due to the large number of objects the scans themselves are expected to take several days.

    We will shutdown the ilifu services at 15:00 SAST tomorrow, 2 September, and we expect ilifu to be offline for several days. Please copy any critical scripts or files you may need to your local machine before the shutdown tomorrow afternoon. You will not have access to your data while the system is offline.

  • The ilifu services and CephFS are again operational.

  • We are pleased to confirm that ilifu services will return online at 12:00 SAST on Wednesday, 16 September 2026, including access to the production CephFS filesystem.

    If you are using the temporary NFS service, please finish your running jobs and log out before 12:00 on Wednesday. Your data on this temporary storage will remain available for a few weeks so that you can transfer it to your usual working area on CephFS. We will communicate the migration deadline and arrangements separately.

    Following extensive recovery and verification work, a small number of metadata inconsistencies remain, affecting a very small fraction of the approximately 1.1 billion files on ilifu. Further recovery attempts for these remaining cases would require keeping ilifu completely offline and unavailable for all users for another 3-4 weeks, with a low probability of success. Having assessed the likely benefit against the impact of extending the outage for all users, we have decided to restore service and address the few outstanding issues separately. For those small numbers of users impacted, we hope you can understand why we have come to this decision. We will be reaching out to those we have identified as being impacted separately, however, if you encounter any file system anomalies, please report them to support@ilifu.ac.za.

    This outage resulted from a highly unusual and unfortunately timed confluence of factors. It has reinforced the need for stronger operational controls, responsible use of shared resources by users and project leads, and additional redundancy measures. We will be reviewing our file-transfer policy and the approval process for research groups, particularly for transfers involving large numbers of files. Storage allocations will be enforced more strictly: once access is restored, please review your usage against your project allocations and delete what you don't need - you'll be improving ilifu's overall resilience in the process.

    We again apologise for the extended outage and recognise the disruption and stress it has caused researchers, particularly students and postdocs approaching thesis or paper submission deadlines. We appreciate your patience and the sustained efforts of the recovery team. Working together, we can strengthen the resilience of ilifu while supporting the diverse needs of our research community. We also look forward to sharing news of new projects and expanded support in the coming months.

  • We are still working on the cephfs metadata recovery, and cephfs existing filesytems have been taken offline. We have bought up new (empty) /users folders in the interim so that people are able to login and work. NB: Old data will not be accessible at the moment while repairs are conducted.

  • 28th August 2026

    CephFS mounted Filesystems CephFS metadata repair requires suspension of ilifu services

    Since the CephFS and MDS issues earlier this week, we have been experiencing issues with CephFS metadata records, a critical component of the filesystem.

    In order to repair the CephFS metadata, CephFS must be taken offline. This means that access to the ilifu cluster will be prevented and all queued and running jobs will be stopped. As this is critical, we will be taking the ilifu cluster down in the next 30 mins. We will restore operations as soon as possible.

  • We are investigating a repeat failure of our cephfs filesystem

  • We have brought ilifu CephFS online again and resumed ilifu operations. ilifu services are now accessible again.

  • The technical team is currently running an analysis on ilifu CephFS in order to determine the next steps for the CephFS metadata recovery. This process may take several hours, and ilifu services continue to be offline during this process.

  • Ilifu services continue to be offline while the technical team works to resolve the critical CephFS metadata issue. We will provide further updates during the day.

  • 24th August 2026

    Ceph Ceph outage

    Ceph / CephFS outage that in the initial investigation seems to be a result of the new updated version of Ceph introduced last week, when running a multiple MDS Ranks.

  • Ilifu operations have been restored. Although users should be able to access ilifu services again, the technical team is still working on the underlying issue.

  • We are currently still experiencing issues with our CephFS instance. Users may experience this as issues accessing ilifu services, accessing file system mounts on compute nodes and general file system performance issues. The technical team is investigating the issue.

  • Systems are back online, please check your slurm jobs as they may have been interrupted by the outage. We are in the process of checking all nodes to ensure they are 100% working with all mounts You may need to restart your Jupyter session.

  • 21st August 2026

    Slurm GPU Nodes GPU nodes allocated to outreach event

    The GPU nodes, gpu-006 and gpu-007, will be unavailable in ilifu Slurm from 17 September - 11 October (previously incorrectly indicated 5 October), as these resources will be dedicated to an outreach event supported by IDIA.

    6th August 2026

    LDAP Issue with authentication service preventing access to ilifu

    We are currently experiencing an issue with a component of the authentication service (LDAP) which is preventing access to ilifu services. The technical team is working to resolve the issue.

  • The LDAP service issue has been resolved and access to ilifu services has been restored.

  • 23rd July 2026

    CephFS mounted Filesystems Filesystem performance issues due to Ceph MDS bug

    We are currently experiencing an issue with the CephFS MDS (metadata server) which is impacting file system performance. Users may experience this as stalled processes or hanging command line prompts. This is predominantly affecting the /idia mount. We are working to resolve this issue. In parallel, we are performing a Ceph minor version upgrade to address the issue.

  • The Ceph minor version update was completed on 2 August. CephFS has completed its recovery process. Users should no longer experience filesystem performance issues.

  • The MDS issue affecting /users has been resolved, users should be able to access services again.

  • We're currently experiencing issues with the CephFS MDS service affecting /users (users' $HOME directories) and access to ilifu services. The technical team is working to resolve the issue.

  • While the primary MDS issue has been resolved, CephFS is still in a recovery state and the upgrade to the latest minor version is still ongoing. Users may still experience filesystem performance issues which may sometimes impact access to ilifu services. The technical team is monitoring the environment and working to bring CephFS into a stable state.

  • We experienced an issue with the CephFS MDS service last night, which resulted in issues with the CephFS mounts, preventing users from logging on to ilifu services. The technical team is currently working to resolve the issue.

  • We are unfortunately still experiencing CephFS MDS issues and are working to resolve this. Users may still experience filesystem performance issues.

  • The MDS issues have been resolved, and access to services have been restored. Users should no longer experience the filesystem performance degradation and hanging processes.

  • The MDS issues are also currently impacting access and performance of the /users, /cbio and /ilifu mounts.