Back to R&D main

V.ISC.2605 - Observability improvements and coverage FY26

Did you know that tuning New Relic alert thresholds and splitting database deadlock signals by application can eliminate false-positive "lost signal" alert storms altogether?

Project start date: 28 September 2025
Project end date: 29 December 2026
Publication date: 31 August 2026
Project status: Completed
Livestock species: All species
Relevant regions: National
Download Report

Summary

Over FY25/26 the New Relic observability platform was tuned for signal quality and coverage. False-positive alert storms were investigated and reduced, and database deadlock alerting was made actionable by separating signals per application and tightening trigger logic. Monitoring was extended to the new LPA and INT environments, with browser and SOAP/API synthetic checks implemented across key login and availability journeys on NLIS, eNVD and MyMLA — including non-production environments aligned to production behaviour. Log ingestion and forwarding were consolidated, targeted logging added for critical scheduled/batch components, and integration endpoints such as eNVD ↔ NLIS brought under 24/7 monitoring.

Inconsistent infrastructure agent versions and failing monthly upgrades were remediated using a validation-first approach in lower environments, and practical documentation on key management and the observability-as-code pipeline was produced and shared with the wider Sysops team. Net result: fewer duplicate alerts, faster triage, better availability confidence and reduced operational risk.

Objectives

Reduce alert noise: Eliminate false-positive alert storms and duplicate notifications so on-call teams only respond to genuine incidents.

Make alerting actionable: Establish clear alert ownership and grouping — per-application deadlock signals, tightened trigger logic — to improve triage speed.
Extend coverage: Onboard new environments (LPA, INT) and critical integration endpoints (eNVD ↔ NLIS) to close monitoring gaps.

Assure critical user journeys: Maintain synthetic coverage of login and availability paths across NLIS, eNVD and MyMLA, in production and non-production.
Improve troubleshooting speed: Consolidate log ingestion and add targeted logging so operational teams can self-serve without manual log pulls.

Standardise platform maintenance: Bring infrastructure agent versions into alignment and enforce a validation-first upgrade path through lower environments.

Build internal capability: Document key management patterns and the observability-as-code pipeline, and transfer knowledge to the wider Sysops team.

Key findings

Alert noise was largely configuration-driven: "Lost signal" style storms traced back to thresholds and time windows rather than genuine service faults — retuning removed the noise once root causes were confirmed.

Deadlock alerts were unusable in aggregate: Bulk signals gave no ownership. Separating them by application and tightening trigger logic converted them into actionable alerts.

Log ingestion was fragmented: Multiple and obsolete server log configurations coexisted, causing confusion over which logs were actually captured.
Critical batch components lacked self-serve logging: Operational teams depended on manual log pulls to troubleshoot scheduled jobs.

Synthetic coverage had blind spots: Non-production environments were unmonitored, and several checks — including SOAP/API synthetics — were failing and eroding trust in uptime reporting.

Agent estate had drifted: Infrastructure agent versions were inconsistent and automated monthly upgrades were not completing reliably, creating an unmanaged patching risk.

Integration endpoints were under-monitored: Links such as eNVD ↔ NLIS carry 24/7 availability expectations but had no matching monitoring coverage.

Benefits to industry

Traceability stays available: Synthetic checks on NLIS, eNVD and MyMLA login and availability journeys mean outages in the national traceability and eDEC systems are detected before producers, agents or processors are blocked.

Fewer disruptions at the saleyard and processor: 24/7 monitoring of the eNVD ↔ NLIS integration protects the movement and consignment data flow that livestock transactions depend on.

Faster resolution: Consolidated logs and actionable alerting cut the time to restore service when something does fail, reducing delays to stock movements and compliance reporting.

Greater confidence in compliance data: Standardised agent maintenance and reduced alert noise lower the risk of silent failures affecting traceability records.

MLA action

Further engage for Fy27

Future research

Yes, we have reengaged this resource in FY27.

More information

Project manager: Gleuto Serafim
Contact email: reports@mla.com.au