L2 Production Support Genesis
Job Req Id:
26989182
Location(s):
Chennai, Tamil Nadu, India
Job Type:
Hybrid
Posted:
Sep. 02, 2026
Discover your future at Citi
Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.
Job Overview
Key Responsibilities
1. Big Data & EAP (Enterprise Application/Analytics Platform) Support
- Support Distributed Environments: Provide Level 2 (L2) and Level 3 (L3) support for applications hosted on Big Data platforms and Citi's Enterprise Application/Analytics Platform (EAP).
- Troubleshoot Data Pipelines: Diagnose and resolve failures in complex data ingestion and processing pipelines, including distributed processing frameworks (e.g., Apache Spark, Hadoop MapReduce).
- Cluster & Resource Monitoring: Monitor cluster resource utilization (using YARN, Cloudera Manager, or similar tools) to identify and resolve memory bottlenecks, queue congestion, and job failures (e.g., Spark Out-Of-Memory errors).
- Data Querying & Validation: Query and validate large-scale datasets stored in distributed data warehouses and file systems (e.g., HDFS, Hive, Impala, or HBase).
- Message Queue Management: Monitor and troubleshoot real-time streaming and messaging platforms (e.g., Apache Kafka), managing consumer groups, offsets, and partition lags.
2. Batch Management & Job Scheduling (Autosys)
- Monitor and manage batch execution: Oversee the execution of critical daily, weekly, and monthly batch processing cycles scheduled via Autosys.
- Troubleshoot batch failures: Rapidly diagnose and resolve Autosys job failures, analyzing log files, identifying dependency issues, and performing necessary job overrides, force-starts, or hold/release actions to minimize business impact.
- Optimize job flows: Collaborate with development and engineering teams to define, configure, and optimize Autosys job definitions using JIL (Job Information Language).
3. Automation & Process Enhancement (Toil Reduction)
- Identify and eliminate manual bottlenecks: Actively analyze daily support activities to identify repetitive, manual tasks ("toil") and design automated solutions to eliminate them.
- Develop automation scripts: Write, test, and deploy robust scripts (using Python, Bash, or PowerShell) to automate routine operations, such as daily health checks, application restarts, log archiving, and data reconciliation.
- Drive process improvements: Evaluate existing support workflows, runbooks, and escalation paths, implementing enhancements to streamline operations and reduce Mean Time to Repair (MTTR).
4. Incident Management & Production Recovery
- Own and drive the end-to-end resolution of L2/L3 production incidents, ensuring strict adherence to corporate Service Level Agreements (SLAs) and Service Level Objectives (SLOs).
- Lead technical triage during Major Incidents (MIM) and high-severity outages. Coordinate effectively with cross-functional global teams (Infrastructure, Database, Networks, Development, and Business Operations) to restore services rapidly.
- Act as the primary technical escalation point during incidents, translating complex technical issues into clear, concise, and business-friendly updates for senior leadership and stakeholders.
- Ensure accurate and timely logging, categorization, and tracking of incidents within ServiceNow.
5. Problem Management & Root Cause Analysis (RCA)
- Lead proactive Problem Management initiatives by analyzing incident trends, identifying systemic patterns, and pinpointing recurring failure points.
- Conduct deep-dive technical investigations—including log analysis, database queries, and infrastructure health checks—to perform comprehensive Root Cause Analysis (RCA).
- Author high-quality Post-Incident Reviews (PIRs) and RCA documents, detailing the timeline, root cause, impact, and preventative actions.
- Collaborate closely with Development and Engineering teams to prioritize, track, and implement permanent bug fixes, structural workarounds, and long-term remediations.
6. Hands-on Unix/Linux & Application Troubleshooting
- Perform deep-dive technical troubleshooting directly within Unix/Linux production environments (analyzing system resources, CPU/memory bottlenecks, process states, and network connectivity).
- Conduct advanced log analysis using Unix command-line utilities (e.g., grep, awk, sed, find, tail) to rapidly isolate application errors and system anomalies.
- Maintain and debug shell scripts (Bash/Korn) used for application startup, shutdown, health checks, and automated maintenance.
7. Database Support & SQL Querying
- Troubleshoot database-related application issues by writing and executing complex SQL queries (including multi-table joins, subqueries, and aggregations) on databases such as Oracle, MS SQL Server, or Sybase.
- Analyze database performance, identify slow-running queries, and collaborate with DBAs to resolve locks, blocks, and indexing issues affecting production.
8. Proactive Monitoring & Alerting with ITRS Geneos
- Actively monitor application health and infrastructure performance using ITRS Geneos (Active Console).
- Configure, customize, and maintain ITRS Geneos samplers, rules, alerts, and Netprobes to ensure comprehensive coverage of critical system components.
Technical Skills & Experience
Core Requirements (Mandatory & Hands-on)
Competency / Technology
Required Hands-on Experience
Big Data & EAP Environments
Working experience supporting Big Data ecosystems and Enterprise Application/Analytics Platforms (EAP). Hands-on familiarity with Hadoop, HDFS, Hive, Spark, YARN, and Kafka. Ability to troubleshoot distributed job failures and monitor cluster health.
Autosys
Strong hands-on experience with Autosys (or similar enterprise job schedulers). Proficient in monitoring batch cycles, troubleshooting job failures, managing dependencies, performing run-time overrides, and writing/modifying JIL (Job Information Language) configurations.
Unix / Linux
Advanced hands-on experience navigating Unix/Linux file systems, managing processes, analyzing system performance (CPU, memory, disk I/O), and writing/debugging Shell scripts (Bash/Ksh).
SQL & Databases
Proficient in writing complex SQL queries to extract, analyze, and troubleshoot data issues across relational databases (Oracle, Sybase, or MS SQL Server). Understanding of database locks, indexing, and basic performance tuning.
ITRS Geneos
Hands-on experience using ITRS Geneos for real-time monitoring. Ability to navigate the Active Console, interpret alerts, and configure basic samplers, rules, and alerts.
Automation & Process Enhancement
Proven experience in automating manual support tasks and optimizing operational processes. Strong ability to identify automation opportunities and implement scripting solutions to reduce operational toil.
Incident & Problem Management
Strong expertise in ITIL processes, specifically Incident and Problem Management. Proven track record of leading major incident triage, managing SLAs, conducting Root Cause Analysis (RCA), and managing the ticket lifecycle in ServiceNow.
Secondary Technical Skills (Highly Desirable)
- Scripting Languages: Strong proficiency in Python or Perl for building automation tools and utility scripts.
- Middleware: Familiarity with middleware technologies (TIBCO EMS, IBM MQ, WebSphere, or Tomcat).
- Cloud & Containers: Basic exposure to containerized environments (OpenShift, Kubernetes) or Cloud platforms (AWS).
- Change & Release Management: Familiarity with ITIL Change Management processes and post-release validation.
Qualifications & Soft Skills
- Education: Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
- Experience: 5+ years of dedicated experience in Application Support or Production Support within a fast-paced financial services or enterprise environment.
- Analytical Thinking: Exceptional problem-solving skills with a methodical approach to diagnosing complex technical issues under pressure.
- Communication: Strong verbal and written communication skills, with the ability to translate complex technical issues into clear business updates for stakeholders.
------------------------------------------------------
Job Family Group:
Technology------------------------------------------------------
Job Family:
Applications Support------------------------------------------------------
Time Type:
Full time------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.
Global Benefits
Discover the top benefits offered to our global workforce, designed to support your well-being, growth and work-life balance. Explore a few of the highlights that make working with us rewarding.
Explore More Jobs
-
대기업 인더스트리얼본부 기업금융심사역
- Seoul, Seoul
-
Wealth Tax Operations Support Senior Supervisor
- Chennai, Tamil Nadu
-
Wealth Relationship Manager - SF Inner Richmond
- San Francisco, California
-
Wealth COO Office - Vice President
- Haryana
-
Early Career Talent Network
Sign up to receive personalized job matches based on your skills and interests. We'll help you discover opportunities that align with your goals.
-
Career Professionals Talent Network
Sign up to receive tailored job matches based on your skills and experience. Discover opportunities that align with your ambitions.