The posting
Summary:
· The successful candidate will be expected to perform a hands-on engineering role, i.e. solving the problem first prior to escalation.
· The candidate is able to read application code, reproduce defects, and land fixes.
· He /she will work alongside the Site Reliability Engineers, who own the longer-horizon work of stopping the same problem recurring.
Must have:
· Minimum 3 years of hands-on experience in resolving production defects in an application support, L2/L3 or run engineering role on systems with real users and real consequences.
· Able to debug someone else's code (Python primarily, TypeScript usefully). Able to take a stack trace and a vague user report and end up at a specific line or configuration value. This is the core skill of the role.
· Able to perform root cause analysis as a discipline. Forming and eliminating hypotheses, stating what is ruled out and why, and distinguishing the cause from the first thing being noticed.
· Possess fluent log and trace reading - Datadog, Splunk, ELK, CloudWatch or similar starting from" a user says it's broken" and getting to a specific failing component.
· Able to perform practical Kubernetes diagnosis. Comfortable with Kubectl: inspecting pods, events, logs and restarts, and recognizing the common failure shapes. Not required to build clusters.
· Experience in scripting using Python or Bash (a must-have requirement). Able to reduce the volume of work and not only absorbs it.
· Clear written skills, both for an RCA a senior engineer will read and for a status update an investment professional will read. Able to handle senior stakeholders with composure and calm.
· A solve-first, escalate-second mindset. Escalation is what happens when the candidate has genuinely exhausted what he/she can do, with investigation attached (hand over a diagnosis instead of hand over a symptom).
· Possess positive learning and collaborative mindset.
· Strong analytical, problem-solving and troubleshooting skills.
· Good written and verbal communication skills.
· Agile, fast learner and able to adapt to changes.
Good to have:
· Experience in provisioning access in an enterprise directory - AD groups, entitlement or approval workflows.
· Deployment or release verification experience. Experience and knowledge in AWS fundamentals - to navigate and understand whatsits where. Experience in building Datadog dashboards and monitors rather than only reading them.
· Experience in supporting an AI or LLM product, where the same input does not always produce the same output and "wrong answer" is a different class of problem from "error".
· Familiarity with databases - SQL and non-SQL.
· Experience in working within a formal change management process.



