TradingView

Senior Site Reliability Engineer

TradingView·LinkedIn·On-site·On-site·3 weeks ago

About the role

TradingView is the world's largest financial analysis platform with more than 100M users across 180+ countries.

We build tools that help traders and investors make informed decisions - from advanced charting and market data to collaboration and publishing features. Our products are used daily by millions of individuals and trusted by companies like Revolut, Binance, and CME Group.

We're continuing to grow and scale our platform, and we're looking for people who care about product quality, take ownership of their work, and want to build systems used by a global audience.

About the team

The team plays a key role in delivering the final product to the production environment and commissioning new services. It works with containerization, automation, orchestration, virtualization, and monitoring systems.

Responsibilities

  • Investigate production incidents and drive them through resolution until all consequences are fully addressed.
  • Perform root cause analysis (RCA) and participate in postmortem reviews together with engineering and product teams.
  • Track and drive corrective actions resulting from incidents and postmortem activities.
  • Develop and improve monitoring, alerting, and observability for assigned services.
  • Define and maintain SLI/SLOs, analyze reliability metrics, and monitor service health objectives.
  • Take ownership of SLA compliance for assigned frontend and backend services.
  • Define and implement availability

requirements

at the service level, in coordination with engineering and product owners.

  • Identify gaps in monitoring and observability and implement improvements to reduce detection and recovery times.
  • Develop and maintain runbooks, troubleshooting guides, and service recovery procedures.
  • Participate in incident validation and escalation reviews.
  • Review and maintain operational and technical documentation.
  • Analyze service performance, resource utilization, and capacity-related risks.
  • Contribute to automation of diagnostics, incident response, and operational workflows.
  • Share opera…
View all