Senior Lead Site Reliability Engineer, Electronic Colo Tradingnew
JPMorgan Chase (JPMC) · Other
- Trading & Execution
- Tokyo-To, Japan
- Professional · Full time
Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.
As a Principal Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms, Electronic Trading Services, you work with your fellow stakeholders to define non-functional requirements (NFRs) and availability targets for the services in your application and product lines. You will be responsible for the reliability, performance, and operational excellence of Linux-based compute platforms that support electronic, colocated (colo) trading. This role will focus on building and operating highly resilient, low-latency infrastructure in data centers where milliseconds matter, with an emphasis on automation, standardization, and disciplined incident management. The ideal candidate combines strong Linux systems administration skills with a production engineering mindset and comfort working close to hardware and networks in a high-availability environment.
Job responsibilities
- Own the build, configuration, and lifecycle management of Linux server fleets supporting colo trading workloads, including provisioning, patching, hardening, and performance tuning.
- Engineer and maintain automation for OS deployment, configuration management, and continuous compliance, with a bias toward reducing manual touch and improving repeatability.
- Partner with network, trading technology, and data center teams to optimize latency, throughput, and stability, including kernel, IRQ, CPU isolation, NUMA, and NIC tuning where appropriate.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
- Operate and improve observability across the stack, including metrics, logs, and alerting, and translate signals into actionable runbooks and service-level improvements.
- Lead incident response for Linux/compute-related events, including rapid triage, mitigation, root-cause analysis, and corrective/preventative actions; drive measurable reduction in recurring incidents.
- Manage hardware-adjacent responsibilities typical of colo environments, including server break/fix coordination, remote hands engagement, firmware alignment, and standardized rack-level practices.
- Implement and maintain secure access patterns, secrets handling, and least-privilege controls aligned to enterprise security and audit expectations.
- Contribute to capacity planning and reliability engineering, including failure-mode thinking, maintenance windows, upgrade strategies, and resiliency testing.
- Produce clear operational documentation, including build standards, runbooks, incident reports, and environment-specific procedures for colo constraints.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Bachelor’s Degree in Computer Science, Cybersecurity, Data Science, or related disciplines
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Professional experience administering Linux in production (e.g., RHEL-derived, Debian/Ubuntu) in a mission-critical environment.
- Strong operational competency in troubleshooting performance and reliability issues across OS, hardware, and basic networking layers.
- Strong expertise in Linux internals, performance tuning, and troubleshooting using enterprise-standard diagnostic tools, with the ability to identify and resolve complex reliability and latency issues in production environments.
- Proficiency in infrastructure automation using Bash and Python or Go, with hands-on experience in configuration management, CI/CD practices, monitoring, observability, and root-cause analysis to drive operational excellence and service reliability.
- Demonstrated experience automating system administration tasks using scripting and/or configuration management.
- Hands-on incident management experience, including participation in on-call rotations and executing structured post-incident reviews.
- Ability to communicate clearly with both engineering and non-engineering stakeholders, including translating technical issues into business impact.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
- Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
- Experience supporting electronic trading, market connectivity platforms, or similarly latency-sensitive, high-availability environments.
- Exposure to colocation data centers, including working with remote hands, controlled maintenance windows, and strict change discipline.
- Experience with performance tuning for low-latency Linux environments (e.g., kernel/CPU pinning strategies, interrupt tuning, time sync discipline).
- Familiarity with infrastructure-as-code patterns and building standardized “golden” server builds.
- Experience operating at scale, including fleet management practices, automation-driven patching, and configuration drift control.