Senior Engineer, Platform Engineering and Architecture
by First Abu Dhabi Bank in Banking & Financial Services
The Senior Engineer, Platform Engineering and Architecture role at First Abu Dhabi Bank (FAB) in Abu Dhabi is responsible for leading the engineering, architecture, and run-state ownership of the bank's enterprise AI & Agentic Platform, ensuring it operates as a reliable, secure, observable, and bank-grade production platform supporting agents and AI workloads across the Group. The role serves as the senior technical authority for platform engineering and architecture across the agentic runtime, model gateway, integration fabric, identity layer, and infrastructure backbone of the platform. It operates in a two-in-a-box model with the existing Platform Product Owner to provide concurrent technical ownership and organizational resilience, with shared accountability for platform availability, performance, cost, security posture, and architectural evolution. The role is deeply technical and requires hands-on engineering depth across distributed systems, agentic protocols, LLM infrastructure, identity and access management, and cloud-native platform engineering. The role owns the end-to-end technical architecture of the AI & Agentic Platform's five-layer stack: Action Gateway, Agent Kernel, Control Plane, Knowledge Foundation, and the Users and Channels layer. It architects and evolves the agentic runtime to support the emerging multi-protocol stack, including Model Context Protocol (MCP) for agent-to-tool access, Agent-to-Agent (A2A) for inter-agent coordination and task delegation, Agent Communication Protocol (ACP) and equivalent emerging standards, while ensuring interoperability with hyperscaler agent fabrics including Azure AI Foundry agents, AWS Bedrock Agents, and Google ADK. It designs the agent kernel including agent lifecycle management, planning and reasoning loops, memory architecture covering short-term, long-term and episodic memory, state management, session affinity, scratchpad persistence, and execution chains. It architects the knowledge foundation including vector store selection and topology, hybrid retrieval using BM25, dense and graph approaches, embeddings strategy, knowledge graph integration using FIBO, OWL, SHACL and SPARQL, context engineering, and grounding patterns. It designs the control plane including policy-gated execution, Know Your Agent (KYA) enforcement at runtime, agent registries, tool registries, capability discovery, evaluation pipelines, and tracing and lineage at agent and tool granularity. The role architects and engineers the platform's LLM gateway and model abstraction layer, providing a unified interface across foundation model providers including Azure AI Foundry, AWS Bedrock, OpenAI, Anthropic, Google Vertex AI and Cohere, with intelligent routing, fallback, retries, prompt and response caching, semantic caching, rate limiting, token accounting, cost attribution and tenant isolation. It designs model serving patterns for managed APIs, dedicated capacity including PTUs / provisioned throughput, and self-hosted open-weight models on GPU infrastructure using vLLM, TGI, Triton or equivalent, balancing cost, latency, sovereignty and compliance. It leads integration architecture between the AI Platform and the bank's core estate, including core banking, payments, treasury, credit and risk systems, the enterprise data platform including Azure, Cloudera and Databricks, enterprise APIs, ESB, event streaming using Kafka and Event Hubs, and the data product layer. It designs the action gateway as the bank's enforcement boundary for agentic action, including API mediation, contract enforcement, circuit breakers, idempotency guarantees, transactional safety and audit-grade action logging. It engineers the platform's tool layer and MCP server estate, including tool packaging, versioning, capability advertisement, schema enforcement and runtime tool discovery across Wholesale, Retail and Group functions. The role architects and owns the agent identity and workload identity model, including non-human identity (NHI) management, agent identity lifecycle, blended user-plus-agent identity for delegated actions and identity propagation across multi-agent flows. It designs authentication and authorization architecture including OAuth 2.1 and OIDC flows for agent-to-tool and agent-to-API interactions, just-in-time credential issuance, short-lived token exchange, mTLS for agent-to-agent communication, and integration with enterprise IAM including Microsoft Entra ID, PAM and secrets management. It implements zero-trust principles across the agentic stack, including least-privilege scoping per agent and per task, real-time policy evaluation, behavioral posture checks and continuous authorization rather than static service-account-style access. It engineers runtime governance controls including KYA enforcement, prompt and output guardrails for PII, PHI and MNPI, prompt injection defense, sensitive action approval flows and human-in-the-loop escalation patterns. It hardens the platform to meet CBUAE, internal model risk and Group governance requirements, including auditability, lineage, data residency, model risk controls, third-party model governance and OWASP LLM Top 10 alignment. The role leads cloud and infrastructure architecture across Azure as the primary platform and AWS, including infrastructure as code using Terraform, networking including private endpoints, peering and egress control, Kubernetes (AKS) and container orchestration, secrets management and CI/CD pipelines. It owns platform Site Reliability Engineering (SRE), including SLO design, error budget management, observability using OpenTelemetry, traces, metrics and logs at agent and tool granularity, incident response, post-mortems, capacity planning and cost governance for a growing fleet of agents and AI workloads in production. The role operates in genuine two-in-a-box with the existing Platform Product Owner, including shared on-call, shared roadmap ownership and shared accountability for major architectural decisions, regulator conversations and critical incidents. It drives engineering excellence through testing discipline covering unit, integration, evaluation and red-teaming, documentation, infrastructure as code maturity, operational readiness reviews and technical mentorship of platform engineers. It represents the platform in senior technical forums with Enterprise Architecture, Cyber, Model Risk, Internal Audit and the Group CTTO's office on architecture and engineering matters and engages with hyperscale, model provider and framework vendor technical teams on platform-level integration, performance, sovereignty and cost optimization. Required technical expertise includes deep current expertise in agentic AI architecture and the modern multi-protocol stack including Model Context Protocol (MCP), Agent-to-Agent (A2A) and Agent Communication Protocol (ACP); hands-on knowledge of agent orchestration frameworks such as LangGraph, Google ADK, LlamaIndex, Autogen, CrewAI and OpenAI Agents SDK; planning and reasoning loops; multi-agent coordination patterns; agent memory systems; evaluation harnesses including RAGAS, OPIK, LangSmith and Promptfoo; LLM serving and inference architecture; managed model APIs; provisioned throughput / PTUs; self-hosted open-weight models; GPU scheduling; vLLM, TGI and Triton; KV-cache management; batching strategies; LLM gateway and AI gateway architecture; model routing; fallback and retry; prompt and semantic caching; rate limiting; tenant isolation; token accounting; policy enforcement; retrieval and knowledge architecture; vector databases including pgvector, Azure AI Search, Pinecone, Weaviate and Qdrant; hybrid retrieval; reranking; embeddings model selection; knowledge graphs including FIBO, RDF, OWL, SHACL, Neo4j and Apache Jena; context engineering; non-human identity (NHI) governance; OAuth 2.1; OIDC; SPIFFE / SPIRE-style workload identity; mTLS; just-in-time credentialing; secret-less architectures; enterprise IAM including Microsoft Entra ID, Azure AD and PAM platforms; cloud-native platform engineering on Azure and AWS; Terraform; Kubernetes (AKS / EKS); Helm; service mesh including Istio and Linkerd; API gateways including APIM, Kong and Envoy; private networking; policy-as-code including OPA and Azure Policy; programming proficiency in Python, Go or Java; SRE; production on-call; incident leadership; post-mortem discipline; SLO and error-budget design; capacity planning; observability tooling including Open Telemetry, Prometheus, Grafana and Datadog; banking integration patterns including Kafka, Event Hubs, API gateways, ESB, ISO 20022, payment rails and core banking integration patterns; security engineering for regulated industries; OWASP LLM Top 10; prompt injection defence; model supply-chain security; secrets management; network segmentation; data residency; audit logging; and production LLM-based or agentic systems. FAB is the largest bank in the UAE and one of the world's largest and safest financial institutions, offering personal and private banking services including credit cards, Islamic banking, investments, loans and mortgages. FAB emphasizes excellence and innovation in providing financial solutions, career development, learning and development initiatives, training and skill development, customer-focused values, and structured recruitment and career progression plans for Emirati talent in the financial and banking sector.