3 días
Expira 21/08/2026
Application Site Reliability Engineer (SRE)
Application Site Reliability Engineer (SRE)
Position Details
- Team: Platform & Production Reliability
- Location: Remote (Americas, LatAm preferred)
- Working Hours: Americas time zones (UTC-3 to UTC-8)
- On-call: Rotation aligned with the London trading day
- Employment Type: Full-time, Permanent
- Experience Level: Mid-Level (3-5 years)
- Technology Stack: .NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform
About The Role
Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.
What You'll Do
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive actions.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views.
- Instrument .NET services to improve telemetry and visibility into service health.
- Define, implement, and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure.
- Improve deployment safety, release automation, and rollback strategies.
- Automate operational tasks through scripting and infrastructure automation.
- Create and maintain runbooks and operational documentation.
- Continuously improve monitoring and platform reliability.
Requirements
Required Technical Skills
- Scripting & Automation
- Observability
- CI/CD & DevOps
- Cloud & Infrastructure
- Databases
- Reliability Engineering
Preferred Qualifications
- Strong experience debugging and supporting .NET/C# applications in production.
- Experience with Windows Server environments and strong PowerShell skills.
- Experience with Grafana, Prometheus, and infrastructure as code tools like Terraform.
- Familiarity with incident response, root cause analysis, and monitoring best practices.
- Exposure to Docker, Kubernetes, or microservices architecture.
Benefits
- Work on mission-critical trading infrastructure that directly impacts customers.
- Solve challenging reliability and scalability problems in a real-time environment.
- Build world-class observability and automation practices.
- Collaborate with experienced engineers in a modern engineering culture.
- Influence reliability strategy and best practices.
If you're passionate about production engineering and building reliable systems at scale, we'd love to hear from you.