Job Description
Senior Site Reliability Engineer
Remote, US
Classy, an affiliate of GoFundMe, is a Public Benefit Corporation and giving platform that enables nonprofits to connect supporters with the causes they care about. Classy’s platform provides powerful and intuitive fundraising tools to convert and retain donors. Since 2011, Classy has helped nonprofits mobilize and empower the world for good by helping them raise over $5 billion. Classy also hosts the Collaborative conference and the Classy Awards to spotlight the innovative work nonprofits are implementing around the globe. For more information, visit www.classy.org.
About the role:
Classy’s Engineering team is hiring a Senior Site Reliability Engineer to join a team of engineers to support and extend our product infrastructure and events management platform to achieve 99.999% availability. The ideal candidate is comfortable leading technical projects and supporting a best-in-class infrastructure and reliability framework. As a member of our team, you will have the opportunity to keep the Classy online fundraising platform for nonprofits running smoothly without interruption.
What you’ll accomplish:
- Play a critical role in contributing to and maintaining a robust, fault-tolerant, global payments platform processing billions of dollars annually.
- Maintain and enhance playbooks and runbooks
- Work across engineering to improve SLO/SLI framework
- Work to improve our incident management process
- Perform daily assessment of the Classy’s stack, distributed public cloud environment for a comprehensive understanding of capacity trending, vulnerabilities, and stability
- Mentor engineers to become proficient developers using best software development practices and processes.
- Participate in an engineering culture of always be learning where the sharing and learning from failures are celebrated and the giving and receiving constructive, candid feedback is highly encouraged.
What you bring (Required):
- Bachelor‘s Degree in Computer Science, or a related field, or equivalent work experience.
- 5+ years building highly scalable projects involving cloud-based infrastructure design and implementation.
- 5+ years of experience using scripting languages such as Python, Bash, or Javascript.
- Strong APM experience using tools such as NewRelic, DataDog, Splunk, etc.
- Solid understanding of microservice architectures, database servers, and API services.
- Solid understanding of distributed data models with experience debugging distributed systems with high data loads.
- High-level proficiency in implementing AWS services (ECS, EKS, EC2, IAM, Cloudwatch, etc.)
- High-level understanding of the application and distributed environments security practice, edge protection observability such as CloudFlare.
- High-level of understanding of infrastructure as code (IaC)
- Experience with Blue/Green deployment infrastructure.
- Ability to understand product requirements and translate them into technical subtasks.
- Experience with Scrum/Agile development methodologies.
- Ability to participate in an on-call rotation for after-hours needs.
What would be awesome to have (preferred):
- Experience building PCI compliant systems
- Experience in payment processing systems
- Experience developing high-volume transaction systems
- Passion for building fault tolerant and secure platforms