Job Description
Title: Senior Site Reliability Engineer – Remote
Location: US
Donnelley Financial Solutions (DFIN) is a leader in risk and compliance solutions, providing insightful technology, industry expertise and data insights to clients across the globe. We’re here to help you make smarter decisions with insightful technology, industry expertise and data insights at every stage of your business and investment lifecycles. As markets fluctuate, regulations evolve and technology advances, we’re there. And through it all, we deliver confidence with the right solutions in moments that matter.
Summary:
We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.
TheSenior Site Reliability Engineeris responsiblefor ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements.
You either have an infrastructure backgroundwitha programmatic, automated mindset or are someone that comes with a software engineering background with infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions.
Responsibilities
- Champion and implement a culture of SRE to maintain a high-quality platform infrastructure
- Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs
- Optimize application performance at scale
- Automate everything including system operational runbooks
- Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies
- Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes
- Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly
- Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations
- Learn continuously and apply lessons learned
- Evangelize best practices, eliminate bottlenecks, and improve process
Qualifications:
- BS in Computer Science or equivalent work experience.
- 2+years experiencewriting software in any modern software language such as C#.NET, Java, Javascript, Node.js, React.
- 2+years experiencecreating automated deployments with tools such as Azure DevOps Pipelines, Ansible, Jenkins or other scripting languages to manage infrastructure, software build and deployment in a continuous integration (CI) environment
- 2+years experiencewriting scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows and Linux environments.
- 2+years experienceimplementing production performance, availability, and scalability monitoring and alerting best practices using a tool such asNew Relic, Dynatrace, DataDog or AppDynamics
- 2+years experienceas a global admin of Azure
- 2+ years of experience supportingpublic client facing revenue generatingsystems
- Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology
- Experience planning, coordinating, developing and executing all stages of test scripts
- Experience securing Windows or Linux systems in 24×7 production environment
- Experience with containerization and managing Kubernetes clusters
- Experience with common networking, firewall and load balancing protocols