Advanced|29 hours|58 lessons

Agentic AI Platform Engineering

Everything an organisation must build so a fleet of agents can run safely in production. The agent runtime and memory architecture, the protocol layer (MCP, A2A), the tool estate, the LLM gateway, the knowledge foundation, agent identity and non-human identity governance, the action enforcement boundary, runtime guardrails, and agent-granular observability. Built for platform engineers and architects who have to operate agents, not write them.

Text-based, no videos
11 modules, 58 lessons
Lifetime access

What you'll learn

Draw the boundary between what an agentic platform owns and what agent teams keep, and defend it in an architecture review
Design the agent runtime: lifecycle, reasoning-loop bounds, the three kinds of memory, session state, and durable execution chains
Support MCP, agent-to-agent coordination, and hyperscaler agent fabrics behind abstractions that survive protocol churn
Run a tool estate as a versioned platform product, with registries, scoped discovery, and schema enforcement at the tool boundary
Build an LLM gateway that gives you provider portability, cost attribution, and enforceable data-residency policy
Operate a shared knowledge foundation: retrieval topology, hybrid search and reranking, embedding migration, and structured grounding
Model agent identity, delegated user-plus-agent authority, and non-human identity governance at population scale
Enforce every real-world action through one unavoidable gateway with idempotency, blast radius limits, and audit-grade logging
Contain prompt injection, sensitive data, and high-risk actions with guardrails and human-in-the-loop approval
Instrument agent-granular observability, evaluation harnesses, red-teaming, and SLOs for probabilistic correctness
Operate the fleet: capacity and cost governance, promotion of non-deterministic versions, incident response, and readiness review
Capstone: design a complete agentic platform for a regulated enterprise, end to end, with explicit tradeoffs

Curriculum

11 modules · 58 lessons
01

The Platform Layer

What an agentic platform is, where the boundary between platform and agent teams sits, how tenancy works for agent workloads, and the reference architecture the rest of the course fills in.

4 lessons
02

The Agent Runtime

Agents as long-lived stateful workloads: lifecycle, reasoning loop controls, the three kinds of memory, session state, durable scratchpads, and supporting orchestration frameworks without being captured by one.

6 lessons
03

The Protocol Layer

Why agent protocols exist, what MCP and agent-to-agent coordination each solve, interoperating with hyperscaler agent fabrics, and designing so protocol churn is a contained change.

5 lessons
04

The Tool Layer

Tools as versioned platform products: the tool contract, registries and breaking-change management, capability discovery as an authorisation decision, schema enforcement, and operating an MCP server estate.

5 lessons
05

The LLM Gateway

The control point above the inference layer: provider abstraction, routing and fallback, caching, token accounting and cost attribution, and policy enforcement at the model boundary.

6 lessons
06

The Knowledge Foundation

Retrieval as shared platform infrastructure: vector store selection and topology, hybrid retrieval and reranking, embedding migration at scale, and structured grounding with knowledge graphs.

5 lessons
07

Agent Identity and Access

The agent-specific identity problem: why service accounts do not fit, governing non-human identity at population scale, delegated user-plus-agent authority, propagation across multi-agent flows, secret-less credentials, and Know Your Agent enforcement at runtime.

7 lessons
08

The Action Gateway

The unavoidable boundary between agent reasoning and real-world effect: API mediation and contract enforcement, idempotency under unknown outcomes, blast radius control, and audit-grade action logging.

5 lessons
09

Runtime Guardrails and Governance

Containment as infrastructure: input and output guardrails, sensitive data detection and routing, prompt injection as an architectural problem, human-in-the-loop approval, and mapping published risk frameworks to real controls.

5 lessons
10

Evaluation and Observability

What agent-granular observability has to capture that conventional tracing does not: reasoning and tool lineage, evaluation harnesses in the promotion path, systematic red-teaming, and SLOs when correctness is probabilistic.

5 lessons
11

Operating the Fleet

Running the platform at fleet scale: capacity and cost governance, promotion of non-deterministic versions, incident response for silent failure, operational readiness review, and a full capstone architecture.

5 lessons

About the Author

Sharon Sahadevan

Sharon Sahadevan

AI Infrastructure Engineer

Building production GPU clusters on Kubernetes. H100s, large-scale model serving, and end-to-end ML infrastructure across Azure and AWS.

10+ years designing cloud-native platforms with deep expertise in Kubernetes orchestration, GitOps (Argo CD), Terraform, and MLOps pipelines for LLM deployment.

Author of KubeNatives, a weekly newsletter read by 3,000+ DevOps and ML engineers for production insights on K8s internals, GPU scheduling, and model-serving patterns.

Ready to master this topic?

Start with the free preview lesson and see for yourself.