Manager -Technology Operation Supportnew
American Express (Amex) · Bank
- Technology & Engineering
- Gurugram, HR, India
- Professional (Band 40 and Below) · Full time
Free job alerts
Get new finance jobs by email.
Fresh vacancies from across banks, asset managers, hedge funds, PE/VC, fintechs and insurers — straight to your inbox. Free, no account, unsubscribe anytime.
The Mgr -Tech Ops Support is a critical role within the Production Assurance Team, responsible for ensuring 24x7 system availability, stability, and operational readiness across infrastructure, application, and network environments. The role acts as the first line of defense for environment health checks, monitoring operations, and ensuring seamless readiness across UAT and Production platforms. The Customer Journey PME team is a cross-functional, collaborative and innovative team responsible for partnering with engineering and product partners to ensure alignment between the organizations and contribute to the key strategic efforts. This strategic role ensures uninterrupted service delivery, rapid incident response, and continuous improvement of operational processes, partnering closely with business and technology teams
Production Assurance Tech Support, Leadership & Governance:
- 24x7 monitoring, incident management, triage, escalation, and resolution activities to maximize service availability and minimize downtime.
- Monitor infrastructure, applications, batch jobs, SFTP, network, database, firewall/security, and end-user computing platforms.
- Act as the shift command lead, orchestrating internal teams and vendor partners to drive timely service restoration and minimize business impact.
- Own end-to-end execution of Incident, Problem, Change, and Release Management processes, ensuring governance, compliance, and SLA adherence.
- Lead major incident bridges, stakeholder communications, escalations, and recovery activities during critical production events.
- Drive accountability across support teams and vendor partners, ensuring corrective actions, RCAs, and problem records are completed and tracked to closure.
- Govern production readiness, operational risk assessments, and change reviews to maintain service stability and resilience.
- Mentor shift analysts and engineers, fostering a culture of ownership, operational excellence, and continuous improvement.
Provide operational insights, risk assessments, and service health reporting to PAT leadership and key stakeholders.
Access, Security & Environment Management:
- Manage access provisioning, SFTP configurations, and credential administration in compliance with security policies.
- Support planned maintenance windows, production releases, and coordinated downtime activities.
Partner with compliance, security, and infrastructure teams to maintain operational controls and platform integrity.
Ticket Management, Reporting & Knowledge Management:
- Manage and govern incident, service request, and problem tickets while ensuring SLA compliance.
- Track, analyze, and report on operational performance, incident trends, and recurring issues.
- Maintain shift handovers, operational logs, knowledge articles, and support documentation.
Provide actionable service metrics, analytics, and operational insights to leadership and stakeholders.
Observability, Reliability, Automation & Continuous Improvement:
- Identify and implement automation opportunities to improve efficiency, reduce manual effort, and enhance operational effectiveness.
- Drive productivity and optimization initiatives within the Production Assurance function.
- Leverage service analytics and operational KPIs to improve reliability, recovery performance, and customer experience.
- Build and enhance observability capabilities through monitoring, alerting, tracing, automation, and self-healing solutions.
Contribute hands-on to enterprise tooling, platform engineering, and reliability initiatives.
Expected Outcomes:
- Deliver industry-leading production stability, availability, and service resilience.
- Reduce recurring incidents through effective problem management and preventive controls.
- Strengthen operational readiness, change governance, and regulatory compliance.
- Enhance disaster recovery preparedness and business continuity execution.
- Provide data-driven insights that enable continuous investment in Production Assurance and platform reliability.
- 8-12 years of experience in Production Support, IT Operations, Application Support, NOC, Command Center, or similar operational support roles.
- Strong understanding of incident management, issue triage, escalation procedures, and service restoration in business-critical environments.
- Experience monitoring applications, infrastructure, batch jobs, SFTP processes, databases, networks, and system health metrics.
- Working knowledge of ITSM processes and ticket management platforms such as ServiceNow, JIRA, Remedy, or equivalent tools.
- Familiarity with change management, release support, disaster recovery testing, and business continuity processes.
- Strong analytical and troubleshooting skills with the ability to quickly diagnose and coordinate resolution of production issues.
- Excellent verbal and written communication skills, with the ability to provide clear and timely updates to technical and business stakeholders.
- Comfortable working in a 24x7 shift-based operating model, including rotational weekends, holidays, and on-call support as needed.
- Good to have Hands-on experience with monitoring and observability tools such as Dynatrace, AppDynamics, Splunk, Grafana, or equivalent platforms.
- Ability to work effectively in a fast-paced, high-pressure operational environment while maintaining attention to detail.
- Must be willing to work from the Gurgaon office a minimum of three days per week.
Preferred Attributes:
- Experience working in Banking, Financial Services, Payments, FinTech, or other highly regulated environments.
- Exposure to command center, mission control, production assurance, or enterprise operations functions.
- Demonstrated experience leading production bridge calls, major incident management, and cross-functional service restoration activities.
- Proven ability to coordinate and influence internal technology teams, business stakeholders, and external vendor partners to achieve desired operational outcomes.
- Strong understanding of Incident, Problem, Change, and Release Management processes, including governance, compliance, and operational risk management.
- Ability to make sound operational decisions under pressure and effectively manage competing priorities during critical incidents.
- Experience leading shift operations, managing escalations, and providing operational direction in a 24x7 support environment.
At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.
As part of Team Amex, you’ll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.
Related finance jobs
- Lead Engineer E4 - Finance Lakehouse
NBS · United Kingdom +2 more
- Domain Architect
SCOR · Bucuresti - Ilfov, Romania
- Credit Support Manager II (AVP)
JPMC · Metro Manila, National Capital Region, Philippines
- Software Engineer II - Adobe Experience Manager, Full Stack
JPMC · Hong Kong
- Software Engineer III - Java AWS
Chase Bank · Bengaluru, Karnataka, India
- Software Engineer III - Java AWS
JPMC · Bengaluru, Karnataka, India
- Software Engineer -II
Socure · chennai, India
- Citco Vilnius Internship Program Autumn 2026 - IT Department
Citco · Vilniaus Apskritis, Lithuania