Systems & Reliability Engineer (SRE / Platform / DevOps)

Systems & Reliability Engineer (SRE / Platform / DevOps)

Full-Time, Bay Area, Hybrid

Location: Bay Area / Hybrid

Type: Full-Time

Level: Senior / Staff

Why this role is exciting

We have built a culture that can ship very quickly. Your job is to make that speed sustainable: releases that are observable and reversible, an on-call process engineers trust, and production systems that surface customer impact early. This role spans release engineering, SRE, platform, production operations, and developer experience. The goal is not to create more process; it is to turn recurring operational work into software and make the safe path the easy path.

About Peer AI

We're a Bay Area-founded, remote-friendly company backed by investors with deep roots in both technology and life sciences. Our platform is live with top-tier pharma customers today, and we're growing fast in a $23B addressable market spanning document authoring, submission workflow management, and agency interactions with the FDA and EMA.

Our Vision

At Peer AI, we are working to clear the path for important scientific and medical discoveries, so that treatments reach the patients who need them as quickly as possible.

Our Values
  • Drive Impact: We focus on delivering real results for our users. Making their work easier, better, and more impactful.

  • Be the Expert: We lead with deep expertise, curiosity, and honesty to guide others toward the best outcomes.

  • Go for Great: We take pride in pushing past “good enough” to deliver standout work in every detail.

  • Win as a Team: We succeed together—building trust, sharing ownership, and helping each other grow every day.

How We Work

Peer AI is an AI-native company, and we want our engineering organization to be AI-native too. We use AI throughout how we build, test, analyze, debug, and operate software. We look for people who ask what can be automated, what should become a system instead of a recurring task, and where human judgment creates the most value. These roles are intentionally broader than their traditional equivalents because we want the people who join us to help redefine the function itself.

Key Responsibilities
  • Own and evolve the path from merge to production, including release coordination, deployment safety, rollback, and release-health signals.

  • Establish and run a lightweight, effective engineering on-call and incident-response process, with clear ownership and useful escalation paths.

  • Improve observability across services, asynchronous jobs, and complete customer workflows, and define reliability signals that reflect what users actually experience.

  • Build and improve CI/CD, automated deployment checks, feature rollout controls, environment reliability, and developer tooling.

  • Use AI to accelerate incident triage, log and trace analysis, root-cause investigation, runbook execution, and post-release analysis where it is genuinely useful.

  • Lead incident reviews focused on systemic fixes, automate recurring remediation, and partner with Quality and Security as we move toward GxP with reproducible change control, release evidence, auditability, and production monitoring.

Must‑Have Qualifications
  • Strong software engineering fundamentals; you should be comfortable solving problems in application code as well as infrastructure.

  • Experience operating distributed production systems and debugging failures across multiple services.

  • Hands-on experience with AWS, containers, CI/CD, infrastructure automation, and modern observability tooling.

  • Experience with production on-call, incident response, deployment strategy, and safe rollback or recovery patterns.

  • A bias toward removing toil through software, with strong judgment about where process adds safety and where it simply adds bureaucracy.

Nice‑to‑Have
  • Experience with AWS ECS/Fargate, SQS, PostgreSQL, CloudWatch, or similar cloud-native stacks.

  • Experience using LLMs or agents for developer productivity, incident response, or production operations.

  • Experience in a regulated, security-sensitive, or otherwise high-reliability software environment.

What We Offer
  • Meaningful equity grant in a high-growth, venture-backed company.

  • Remote-first setup with quarterly on-sites in San Francisco.

  • Comprehensive health coverage, 12 weeks paid parental leave, and monthly home-office stipends.

  • Unlimited PTO and a mission that directly accelerates life-saving clinical research.

How to Apply

Send your resume to careers@getpeer.ai with the subject line: Systems & Reliability Engineer, [Your Name]. A short note is welcome: tell us about an operational or release process you turned into a system - especially one that made engineers faster and production safer at the same time.

Equal Opportunity

Peer AI is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. If you need a reasonable accommodation during the hiring process, please let us know.

Hiring Manager

Head of Engineering