Search jobs > San Francisco, CA > Senior site reliability

Senior Site Reliability Engineer

Grindr
Chicago; Palo Alto; San Francisco
Full-time

This is a hybrid role based in our Chicago, Palo Alto or San Francisco office and will require you to be in office Tuesdays and Thursdays.

What’s so interesting about this role?

As we enter our second year as a public company, Grindr is building on the success we’ve had over our 15-year history in connecting, supporting, and improving the lives of the LGBTQ+ community globally.

We are hiring a Site Reliability Engineer to join our newly established SRE team. You will work closely with our cloud engineering and software development teams to design, implement, and maintain systems that ensure the high availability, performance, and security of our platform.

This is a unique opportunity to shape the SRE culture and practices from the ground up, influencing the way we deliver and manage our services.

What’s the job?

Monitoring and Alerting : Set up and maintain monitoring systems to track the health and performance of applications and infrastructure.

Create and manage alerting mechanisms to detect and respond to issues quickly.

  • Incident Response : Handle incidents and outages, working to resolve them swiftly and minimize downtime. Performing root cause analysis to prevent future occurrences and improve system resilience.
  • Automation : Develop tools and scripts to automate repetitive tasks, such as deployments, monitoring, and scaling, to increase efficiency and reduce human error.
  • Performance Optimization : Analyze system performance and identify bottlenecks or areas for improvement. Work with development teams to optimize code and infrastructure for better performance and resource utilization.
  • Capacity Planning : Plan for future growth by analyzing current usage trends and forecasting resource needs. Additionally, you’ll ensure that systems can handle increased load without compromising performance or reliability.
  • Service Level Objectives (SLOs) and Service Level Agreements (SLAs) : Define and measure SLOs and SLAs to set expectations for system reliability and performance.

Track these metrics and work to maintain or exceed the defined standards.

Incident Management and Postmortems : After incidents, conduct post mortems to document what went wrong, what was done to fix it, and how to prevent similar incidents in the future.

This process helps in continuous improvement and learning from failures.

Collaboration with Development Teams : Work closely with software developers to integrate reliability and performance into the development process.

Provide guidance on best practices and assist with designing resilient systems.

  • Security and Compliance : Ensure that systems are secure and compliant with relevant regulations and standards. They implement security measures, monitor for vulnerabilities, and respond to security incidents.
  • Continuous Improvement : Continuously look for ways to improve system reliability, performance, and efficiency. Stay updated with industry trends and advancements to implement the best practices and technologies.
  • Participate in an on-call rotation

What we'll love about you :

  • Technical Expertise :
  • Proficient in at least one programming language (e.g., Python, Go, Java).
  • Strong knowledge of Linux / Unix systems.
  • Experience with cloud platforms (e.g., AWS, GCP, Azure).
  • Familiarity with containerization and orchestration tools (e.g., Docker, Kubernetes).
  • Understanding of networking concepts and protocols.
  • Reliability Engineering :
  • Experience with monitoring, logging, and alerting tools (e.g., Prometheus, Grafana, ELK stack).
  • Ability to implement and manage CI / CD pipelines.
  • Knowledge of infrastructure as code (e.g., Terraform, Ansible).
  • Proficiency in automated testing and deployment practices.
  • Understanding of SRE principles and practices, including SLAs, SLOs, and SLIs.
  • Security :
  • Knowledge of security best practices and compliance standards.
  • Experience with vulnerability assessment and mitigation.
  • Operational Excellence :
  • Proven track record of maintaining high availability and performance in production environments.
  • Experience with incident management and post-mortem analysis.
  • Ability to optimize system performance and resource utilization.

Basic Qualifications :

  • 5+ years of experience in site reliability including incident response, incident management, automation and performance optimization
  • 5+ years of experience in cloud platforms (AWS preferred)
  • 4+ years of experience working with DevOps technologies such as Docker, Kubernetes, Helm, and Terraform
  • 4+ years developing and maintaining CI / CD pipelines
  • 4+ years experience using a scripting language like python or bash
  • Experience coding in Kotlin or another JVM language is a plus

What you'll love about us

Mission and Impact : Grindr is building the global gayborhood in your pocket. Your role will impact the lives of millions of LGBTQ+ people around the world.

Through our success, we are making a world where the lives of our community are free, equal, and just.

  • Family Insurance : Insurance premium coverage for health, dental, and vision for you and partial coverage for your dependents.
  • Retirement Savings : Generous 401K plan with 6% match and immediate vest in the U.S.
  • Compensation : Industry-competitive compensation and eligibility for company bonus and equity programs.
  • Queer-Inclusive Benefits : Industry-leading gender-affirming offerings with up to 90% cost coverage, access to Included Health, monthly stipends for HRT, and more.
  • Additional Benefits : Flexible vacation policy, monthly stipends for cell phone, internet, wellness, food, and commuting, breakfast / lunch provided onsite, and yearly travel & leisure stipend.
  • 30+ days ago
Related jobs
Promoted
Swish Analytics
San Francisco, California
Remote

The Swish Analytics DevSecOps and Infrastructure team is looking for an experienced Site Reliability Engineer who will support our enterprise infrastructure. We believe that oddsmaking is a challenge rooted in engineering, mathematics, and sports betting expertise; not intuition. Work closely with t...

Promoted
Outdefine
San Francisco, California

Read the overview of this opportunity to understand what skills, including and relevant soft skills and software package proficiencies, are required.Long term contract: 1099 – US Based candidates / No C2C or C2H.Demonstrated ability to deliver solutions that are easily maintainable, understandable, ...

Promoted
Cisco Systems
San Francisco, California

As a Principal Site Reliability you will focus on innovating and providing strong technical vision as well as work with the team to build reliable, scalable and highly available datastores on a constantly growing multi-region scale platform. We’re looking for a reliability-focused engineering leader...

Promoted
Retool Inc.
San Francisco, California

As our first Site Reliability Engineer, you will be instrumental in defining and shaping the processes and practices for a pivotal new business offering. This role requires a blend of deep technical expertise in site reliability engineering and a keen product sense to create solutions that not only ...

Promoted
Xero
San Francisco, California

As a Senior Software Engineer in Data Reliability Engineering, you will architect and build platform products that reduce cognitive load, abstract complexity, and create the necessary controls for safe database ownership at Xero. At Xero, our Data Reliability Engineering team plays a critical role i...

Promoted
Patreon
San Francisco, California

Advocate and implement Site Reliability Engineering practices including SLIs, SLOs, and SLAs across the engineering organization to improve our operational excellence. You have experience in DevOps, Site Reliability, or backend/infrastructure engineering for a company experiencing fast-paced growth....

CIRCLE
San Francisco, California

As a Senior Site Reliability Engineer at Circle, you will design, build, and maintain Circle’s infrastructure estate to meet the growing worldwide customer base on public cloud providers across multiple regions. Senior Site Reliability Engineer (III). Senior Site Reliability Engineer (III). All the ...

AutoRABIT Holding Inc.
San Francisco, California

About the role: AutoRABIT is looking for a Senior Site Reliability/DevSecOps Engineer to help develop, scale and operate our cloud services  In this role you will be an experienced business professional able to implement and execute best practice operations and improvements across teams by prov...

Cisco Meraki
San Francisco, California
Remote

Building an automatic service lifecycle platform to coordinate the full lifecycle of all infrastructure (server, storage, network and site). Deploying comprehensive monitoring tools to provide insight into the performance and reliability of our infrastructure. ...

Federal Reserve System
San Francisco, California

Site Reliability Engineer, you will be part of the Data & Analytics Services (DAS) Team and will get an opportunity to broadly apply your engineering skills across various technology solutions, as well as build your skills in other areas by being exposed to various aspects of product delivery from i...