Job offer
Senior Principal Infrastructure Services (SRE Practice)
Northern Trust is seeking an experienced Senior Principal Java Reliability Engineer for its SRE practice in Austin, TX, to develop automation solutions for complex cloud and hybrid infrastructures. The role involves leading system architecture and incident management, as well as mentoring teams to foster a culture of resilience and continuous improvement.
Tasks
- Develop a deep understanding of Northern Trust’s business processes and services, and effectively communicate the impact of technology outages to senior management, non-technical stakeholders, and distributed teams.
- Provide a robust, scalable, and highly available SRE automation foundation that spans multiple service platforms, including Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), hybrid on-premises and colocation systems, Microsoft Windows, and Linux/Unix-based platforms.
- Conduct root-cause analyses, translate complex technical issues into understandable business impacts, develop comprehensive remediation strategies, establish preventive measures, and provide continuous monitoring services.
- Lead the design and further development of highly resilient, scalable, and high-performance distributed systems that span numerous services and application areas.
- Collaborate with product owners and application teams to influence system designs that improve observability, automation, and operational agility; positively impact team culture; build resilient systems; and effectively manage external load shifts.
- Define and promote reliability patterns, architectural and technical practices, and tool configurations that align with business outcomes.
- Drive an automation strategy by designing and developing tests, tools, and pipelines that reduce manual work, optimize workflows, and improve both the customer experience and the employee experience.
- Design an evolving engineering process that encompasses infrastructure, configuration, and software deployments, incorporating automated operational controls and security measures, including embedded, non-disruptive operational workflows.
- Participate in or conduct root-cause analyses for production incidents to ensure timely mitigation and resolution during the recovery phase.
- Identify and implement capacity metrics, risks, and threats for one-time or ongoing reliability systems to enable regular assessments and prioritization of services.
- Architecturally design and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Alerting and Escalation Response Teams (ARTERs) that provide clear reports on service health, performance, and risks.
- Establish metrics and targets for SLIs, SLOs, and error budgets that reflect customer-centric service-level expectations.
- Initiate efforts that use case-by-case analysis, recent changes, regression testing, and other factors to visualize system failure modes.
- Communicate industry-standard operational best practices, measure them (SRE), turn your vision into a better SRE team, and ensure service reliability.
- Create and maintain team, codebase, and knowledge documentation using industry-standard architectures, standards, and patterns, as well as incident postmortems.
- Proactively share service capacity metrics, availability, performance, and fault tolerance with peer teams.
- Work closely with product, development, platform, security, and operations teams to integrate SRE principles into team frameworks.
- Act as a trusted advisor, mentoring compliance teams and operational staff to shape business insights for the benefit of Northern Trust and its external stakeholders.
- Drive the adoption of strategic SRE programs that measurably improve cross-functional collaboration, incident response, and operational efficiency.
- Manage and prioritize multiple workload initiatives (including technical and business system operations needs, with a focus on cross-functional collaboration).
Job details