Minor Service Outage

Setonix Partial Outage
Login nodes Operational
Data-mover nodes Operational
Slurm scheduler Operational
Setonix work partition Operational
Setonix debug partition Operational
Setonix long partition Operational
Setonix copy partition Operational
Setonix askaprt partition Operational
Setonix highmem partition Operational
Setonix gpu partition Operational
Setonix gpu high mem partition Operational
Setonix gpu debug partition Partial Outage
Lustre filesystems Operational
/scratch filesystem Operational
/software filesystem Operational
/askapbuffer filesystem Operational
/askapingest filesystem Operational
Storage Systems Operational
Acacia Ingest Operational
Acacia MWA Operational
Acacia Projects Operational
Banksia Operational
Data Portal Systems Operational
MWA ASVO Operational
ASKAP Operational
ASKAP ingest nodes Operational
ASKAP service nodes Operational
Central Services Operational
Authentication and Authorization Operational
Service Desk Operational
License Server Operational
Application Portal Operational
Origin Operational
/home filesystem Operational
/pawsey filesystem Operational
Central Slurm Database Operational
Documentation Operational
Visualisation Services Operational
Remote Vis Operational
Vis scheduler Operational
Setonix vis nodes Operational
Nebula vis nodes Operational
Visualisation Lab Operational
Reservation Operational
CARTA - Stable Operational
CARTA - Test Operational
Pawsey Remote VR Operational
The Australian Biocommons Operational
Fgenesh++ Operational
Operational
Degraded Performance
Partial Outage
Major Outage
Maintenance
Allocated Cores (Setonix)
Fetching
Allocated Nodes (Setonix work partition)
Fetching
Allocated nodes (Setonix askaprt partition) ?
Fetching

Sep 7, 2026

Completed - All services were returned to production on Thursday (3rd September 2026).

The Slurm controller on Setonix has been operating normally over the weekend.

With the Slurm upgrade there are changes in how Slurm allocates resources for jobs running on GPU nodes. As a result, researchers need to update their job submission scripts to ensure continued compatibility. Please read through the documentation page prepared by Pawsey staff (https://pawsey.atlassian.net/wiki/spaces/US/pages/1965031425/SLURM+25.11.7+changes).

If you have any questions, please reach out to the Service Desk (help@pawsey.org.au).

Thank you for your patience. And thank you to the Pawsey staff who resolved the issue.

Sep 7, 09:22 AWST
Update - We are continuing to verify the maintenance items.
Sep 3, 11:26 AWST
Update - Pawsey staff have tweaked the configuration of the Slurm controller for Setonix and brought more of the system back online. We will be providing NVIDIA additional logs and will continue to monitor the system.
Sep 3, 10:53 AWST
Update - We are continuing to investigate the issue.
Sep 2, 09:17 AWST
Update - The Pawsey team have altered the network connections on the scheduler to make it more responsive.

A ticket is being lodged with NVIDIA, who provide L3 Slurm support.

Sep 1, 19:12 AWST
Update - Setonix passed our internal testing, but once the maintenance reservation was removed there appeared to be communication issues between the Slurm controller and nodes.

We are investigating the issue.

Sep 1, 17:37 AWST
Verifying - Vis services are back up in Production.

The Setonix workflow nodes are being looked at.

Windows updates are being applied to the Nebula nodes.

Sep 1, 16:09 AWST
Update - Acacia Projects is back up in Production.

Acacia MWA is back up in Production.

Core services have been patched.

Sep 1, 15:49 AWST
Update - Acacia Ingest is back up in Production.

Banksia is back up in Production.

Mediaflux is back up in Production.

Sep 1, 14:34 AWST
In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Sep 1, 08:00 AWST
Scheduled - Maintenance on Pawsey systems will be carried out on Tuesday the 1st September to apply required patches and updates to improve the systems stability, security, and performance. This maintenance window will also be used to undertake other tasks which require down-time to achieve.

Planned work for this window includes:
• Patching of visualisation services will be undertaken
• Mediaflux server will have the latest bug and security fixes from Rocky Linux
• Finalisation of the control plane of Acacia Ingest
• Acacia compliance runs
• Banksia's servers will have the latest bug and security fixes from Rocky Linux
• Banksia will be upgraded to the latest version of ScoutAM
• Patching of core Pawsey services will be undertaken
• Setonix will have the latest bug and security fixes from SLES 15 applied
• Setonix will be upgraded to a supported version of Slurm
• ASKAP Ingest will have the latest bug and security fixes from SLES 15 applied
• ASKAP Ingest will be upgraded to a supported version of Slurm

With the Slurm upgrade there are changes in how Slurm allocates resources for jobs running on GPU nodes. As a result, researchers need to update their job submission scripts to ensure continued compatibility. Please read through the documentation page prepared by Pawsey staff (https://pawsey.atlassian.net/wiki/spaces/US/pages/1965031425/SLURM+25.11.7+changes).

If you have any questions, please contact help@pawsey.org.au.

Aug 25, 09:00 AWST

Sep 6, 2026

No incidents reported.

Sep 5, 2026

No incidents reported.

Sep 4, 2026

No incidents reported.

Sep 3, 2026

Sep 2, 2026

Sep 1, 2026

Aug 31, 2026

No incidents reported.

Aug 30, 2026

No incidents reported.

Aug 29, 2026

No incidents reported.

Aug 28, 2026

No incidents reported.

Aug 27, 2026

No incidents reported.

Aug 26, 2026

No incidents reported.

Aug 25, 2026

No incidents reported.

Aug 24, 2026

No incidents reported.