About the Role
Sr. Systems Site Reliability Engineer
Full time
job requisition id
JR466324
Who is Forcepoint?
Forcepoint simplifies security for global businesses and governments. Forcepoint’s all-in-one, truly cloud-native platform makes it easy to adopt Zero Trust and prevent the theft or loss of sensitive data and intellectual property no matter where people are working. 20+ years in business. 2.7k employees. 150 countries. 11k+ customers. 300+ patents. If our mission excites you, you’re in the right place; we want you to bring your own energy to help us create a safer world. All we’re missing is you!
The ideal candidate will be Dallas based (for a hybrid role) & have a broad background spanning both applications and infrastructure. They will have direct experience with multiple coding language, core SRE practices & methodologies.
Open to US Remote
Essential Functions
Solve problems relating to mission critical services and build automation to prevent problem recurrence, with the goal of automating response to all non-exceptional service conditions. This individual will be focused on maximum availability, reliability, security, and performance for Forcepoint services.
Requirements:
- Designs and Maintains secure, scalable, and highly available architectures for On-Prem and Cloud Hosted environments
- Fully understands Agile Systems Engineering practices
- Strong understanding of cloud-based architecture and cloud operations. Hands-on experience with Amazon Web Services and/or equivalent public cloud technology
- Experience in administration/build/management of Linux systems
- Foundational understanding of Infrastructure and Platform Technology stacks
- Strong understanding of Networking concepts and theories, such as different protocols (TCP/IP, UDP, routing protocols, etc), VLAN configuration, DNS, OSI layers, and load balancing
- Understanding of security architecture and certificate management
- Working knowledge of Infrastructure and Application monitoring platforms such as Grafana Cloud, Solarwinds, NewRelic, DataDog etc.
- Working knowledge of Incident Response and Alerting platforms such as PagerDuty, Opsgenie, XMatters etc.
- Understanding of the core DevOps practices (CI/CD pipeline, release management etc.)
- Ability to write code using any one modern programming language (Python, JavaScript, Ruby etc.). Additional scripting skills are preferred
- Configuration management platform understanding and experience (Chef/Puppet/Ansible)
- Prior experience in Cloud management automation tools (Terraform/CloudFormation etc.) is crucial
- Experience with source code management software and API automation is crucial
- Cloud certifications or equivalent experience is highly regarded
- Service availability oriented mindset with a pro-active approach to problem solving. An ideal candidate should be able to develop automated solutions to prevent recurring problems
- Possesses the ability and willingness to challenge the status-quo and optimize current procedures and processes
- Creates flowcharts, diagrams, and other documentation
- Knowledge of Container tools Docker/ Kubernetes,
- Benchmarks applications and services performance and design scalable systems and APIs
- Must have strong Linux experience supporting production systems
- Additional Qualification:
- Expertise in designing, analyzing and troubleshooting large-scale distributed systems.
- Understanding of Unix/Linux systems from kernel to shell and beyond, taking in system libraries, file systems, and client-server protocols along the way
- Good knowledge of virtualization technologies and container technologies
- Experience with containers and HA clusters; experience with Docker and Amazon ECS /Kubernetes/ Mesosphere/Docker Swarm a plus VMware certification is preferred.
- Proficient Knowledge or Application of Agile/Scaled Agile: Familiarity with agile methodologies (such as Scrum or Lean) and experience in applying them to IT development or project management. Understanding of scaled agile frameworks (SAFe) is a plus.