About the Role

Title: Site Reliability Engineer (global, remote)

Location: Boston MA US

Job Description:

PacketFabric

Site Reliability Engineer

THE ROLE:

Our SRE team is a small and close-knit group of experts spread across the world, and we are looking to expand our team in Europe and NA. Our role is to help to make business and development operations continue to flow smoothly, as well as building on our existing processes and platforms to reduce costs, increase stability, and enable the rest of the business to have as friction-free experience as possible.

WHAT YOULL BE DOING:

  • Troubleshooting issues along with developers, providing systems level and architecture insight to the current issue.
  • Assisting and consulting with developers, network engineers, etc to ensure our software and platforms adhere to best practices and provide the best outcomes
  • Working with our configuration management systems to improve automation workflows and support requested features
  • Maintaining and driving forward our compute infrastructure, comprising of both cloud resources and bare metal devices
  • Solving complex and/or unintuitive system stability issues.
  • Researching, investigating, and providing justification for new technologies that would benefit the organisation

WHAT YOU BRING:

As a well-rounded systems engineer and automation enjoyer with a diverse set of skills, this makes you one of the very best people to troubleshoot, monitor the platform, and be on top of releases. You need:

  • A high level of Linux systems knowledge and experience
  • The ability to identify risk and manage appropriately
  • The ability to work autonomously as well as with colleagues as required
  • The ability and willingness to learn new things, which may require figuring out without documentation – and then write documentation afterwards!
  • A high degree of drive to improve and automate your environment with minimal guidance
  • Experience with automation of cloud platform (GCP, AWS, etc) management
  • Extensive experience with Docker and Kubernetes as well as other forms of containerisation and virtualisation
  • Experience working with Python – coding as well as managing application/package deployments
  • Experience with message queue systems like RMQ, ZMQ, Kafka, Pub/Sub etc
  • Experience with configuration management/IaC tools such as Ansible, Salt, Terraform, Chef, Puppet, CloudFormation etc
  • Experience with CI/CD pipelines using Jenkins, GitHub Actions, BitBucket, Bamboo etc
  • Solid understanding of TCP/IP, including knowledge of common protocols such as HTTP, TLS, DNS, DHCP, NTP, SSH, SMTP, etc
  • Solid understanding of nginx and SSL

PREFERRED EXPERIENCE:

  • Familiarity with network platforms (Juniper, Cisco, Arista, Ciena, Nokia, etc)
  • Familiarity with networking concepts such as VLAN, VXLAN, MPLS, BGP, etc
  • Experience with large scale network management and/or monitoring.
  • Hands-on experience making applications work at scale.
  • RDBMS experience, preferably PostgreSQL
  • Experience with time-series data stores
  • Experience in PXE based deployments
  • Experience working in an environment leveraging remote communication collaboration tools like slack, zoom etc. across multiple time zones
  • Knowledge of server hardware
  • Larger-scale software development experience, ideally with Python
  • Experience with multiple programming languages

APPLY HERE