finseats

Associate, Site Reliability Engineer (Platform Support), SRE and Governance, Group Technology

DBS Bank (DBS Bank) · Bank

Responsibilities:

  • Administrate the full Kubernetes platform life cycle to ensure the platform remains secure, reliable and highly available.
  • Work with other engineers to support the infrastructure of OpenShift as well as assist to resolve infrastructure queries/issues of OpenShift tenant applications.
  • Automation of platform operations and management.
  • Configure and Manage Monitoring, alerts for the OpenShift environments.
  • Manage capacity for Kubernetes environments.
  • Hands-on experience with Docker containerization, Kubernetes orchestration and tools such as Jira and Jenkins.
  • Hands-on experience in handling Linux OS such as Patching, Filesystem management, User management, SELINUX etc.
  • Administration and support Windows Servers and Linux server environments.
  • Perform server hardening, configuration, patching, upgrades, and maintenance.
  • Troubleshoot IIS, SSL certificate, OpenSSH, tectia and IBM CD file transfer
  • Monitor server health, CPU, memory, disk, services, and system availability.
  • Troubleshoot OS, application, network-connectivity, and performance issues.
  • Handle Linux services, packages, file permissions, SSH, cron jobs, and system logs.
  • Respond to incidents, alerts, service requests, and production changes.
  • Perform vulnerability remediation and OS hardening.
  • Maintain operational documentation, SOPs, and troubleshooting guides.

     

Requirements:

  • 3 years of work experience with a bachelor’s degree in computer science or related
    Preferred Qualifications. At least 1 year’s hands-on experience with containers in Production Environments - Docker, OpenShift, Kubernetes, Linux and Window Servers preferred
  • Provide operational support for OpenShift Container Platform (OCP), Red Hat Linux, Windows Server, and infrastructure platforms across production and non-production environments.
  • Deliver 24x7 operational support for infrastructure and platform services, ensuring timely incident resolution and service restoration.
  • Perform system administration, monitoring, troubleshooting, performance tuning, and capacity management for Linux, Windows, and containerized environments.
  • Participate in incident management, problem management, root cause analysis (RCA), and post-incident reviews to improve operational stability.
  • Maintain operational documentation, standard operating procedures (SOPs), knowledge articles, and support runbooks.
  • Support security and compliance requirements by executing vulnerability remediation, patch management, access control, and platform hardening activities.
  • Experience with configuration management tools (Chef, Ansible, terraform etc.).
  • Hardening, securing the Kubernetes cluster with monitoring and auditing dashboards
  • Knowledge in infrastructure technologies such as HP and DELL hardware (Blades and Rack servers)
  • Excellent verbal, written, skills; in particular, demonstrated ability to effectively communicate technical and business issues and solutions to multiple organizational levels internally and externally.
  • Candidate must have demonstrated and be prepared to exhibit initiative and ownership of consistent delivery success
  • Be scheduled On-Call to support the infrastructure and systems

Location:

DBS Asia Hub

Job:

Technology

Schedule:

Regular

Employee Status:

Full time

Free job alerts

Get more Technology & Engineering jobs like this.

Fresh vacancies from across banks, asset managers, hedge funds, PE/VC, fintechs and insurers — straight to your inbox. Free, no account, unsubscribe anytime.

We only store your email + filters · privacy