PowerPlan, Inc

Principal Cloud Ops Engineer

Job Locations US-GA-Atlanta
Posted Date 2 weeks ago(8/28/2026 3:44 PM)
Job ID
2026-1972
# of Openings
1
Category
Information Technology

Overview

PowerPlan is seeking a Principal Site Reliability Engineer to sit at the heart of our cloud platform's reliability, scalability, and operational maturity. You'll work hands-on across AWS and Azure environments, solving complex production problems while systematically eliminating the manual toil that creates them.

 

This role offers significant autonomy, deep technical impact, and the opportunity to shape how reliability engineering is practiced across the organization.

About PowerPlan

PowerPlan helps the companies that power the world unlock greater value from their infrastructure investments. We combine deep industry expertise, trusted technology, and AI-driven innovation to help capital-intensive organizations manage the financial complexity of their assets with confidence.

 

We're in the middle of one of the most significant transformations in our history. Decades of proven functionality are being reimagined on a modern SaaS foundation, while AI becomes central to how we build, deliver, and evolve our products. The next generation of the PowerPlan platform is being designed right now. Join us, and you'll have the opportunity to influence the technology, engineering practices, and capabilities that will shape it for years to come.

Responsibilities

Your Impact

  • Resolve escalated infrastructure cases across AWS and Azure and deliver targeted automations that reduce manual resolution time.
  • Eliminate or significantly reduce manual intervention for the highest-frequency operational issues through automation and tooling.
  • Establish a consistent, high-quality incident response and post-incident review process for critical production incidents.
  • Deliver a mature, SLO-aligned observability platform with dashboards, tuned alerts, and clear reporting.
  • Coach teams on effective incident communication and decision-making.

What Success Looks Like

Within 90 days, you'll have shipped automations that measurably cut manual resolution time on recurring issues. By month six, you'll have eliminated toil on the highest-frequency operational problems. By month twelve, on-call and engineering teams will be running on an observability platform you built — one with dashboards, tuned alerts, and SLIs/SLOs that make reliability a data-driven practice instead of a guessing game.

Qualifications

What You'll Bring

  • Deep hands-on experience operating production systems in AWS and Azure environments.
  • Strong automation skills using Python and PowerShell in operational contexts.
  • Proven ability to identify repetitive operational work and eliminate it through automation.
  • Experience leading incident response and blameless post-incident reviews.
  • Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring.
  • Ability to influence engineering practices without formal authority.
  • Clear written and verbal communication skills across technical and non-technical audiences.

Education & Experience

  • Extensive experience in cloud operations, site reliability engineering, or infrastructure engineering roles, or equivalent professional experience.

 

 

PowerPlan is an EOE

Applicant and Candidate Privacy Notice

 

Please note that this is a hybrid role that involves a combination of onsite work from our corporate office as well as work from home. While we strive to accommodate flexible working arrangements when sensible, there will be times when onsite work is required. This could include scheduled office days, team meetings, client meetings, or special events.

Options

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.
Share on your newsfeed