Agentic Engineering Roadmap

One track from three prerequisites taught elsewhere to production agent systems; start at LLM Fundamentals. Every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.

Where to start

Agentic engineering

19 stages · 0/73 lessons

Context, tools, retrieval and the agent loop, then the evals, tracing, guardrails and reliability that take an agent to production.

  1. Software Engineering
  2. Python / TypeScript
  3. APIs & Async Programming
  4. LLM Fundamentals
  5. Prompt & Context Engineering
  6. Structured Outputs
  7. Tool Calling
  8. RAG
  9. Agent Loops
  10. State & Memory
  11. MCP
  12. Workflow Orchestration
  13. Multi-Agent Systems
  14. Evaluation
  15. Observability
  16. Guardrails & Security
  17. Human-in-the-Loop
  18. Reliability
  19. Production Agent Systems
0 / 73 lessons masteredNot started 73Learning 0Practicing 0Mastered 0
  1. 1

    Software Engineering

    Start hereTaught elsewhere

    Agents are ordinary distributed systems with a probabilistic component in the middle, so the ordinary discipline applies first: clear interfaces, deliberate error handling, tests, and reasoning about state and failure. This stage is taught in the Backend roadmap; the agentic lessons assume it.

    Before moving on: Design, test and debug a service with clear interfaces, handle errors deliberately, and reason about state, concurrency and failure.

    Backend roadmap →
  2. 2

    Python / TypeScript

    Taught elsewhere

    The ecosystem is built in two languages, and every example on this site is written in one of them: typed models (Pydantic / Zod), async / await, packaging, and reading library source when the docs run out. Every DSA topic ships a Python and a TypeScript implementation, which is a good place to practise.

    Before moving on: Read and write idiomatic, typed, async Python or TypeScript and follow a library into its source when the documentation stops.

    DSA roadmap →
  3. 3

    APIs & Async Programming

    Taught elsewhere

    Every model call, tool call and retrieval is a network request. Authentication, streaming, timeouts, retries with exponential backoff and rate limits are the vocabulary the tool-calling and reliability stages build on; the API Design roadmap teaches that side of the contract. Running independent calls concurrently without corrupting shared state is how an agent stays fast, and that half lives in the Concurrency domain — start with Await Is a Yield Point, The Sequential Await Trap and Shared Mutable State.

    Before moving on: Call an HTTP/JSON API with auth, streaming, timeouts, backoff and rate limiting, and run independent calls concurrently without corrupting shared state.

    API Design roadmap →
  4. 4

    LLM Fundamentals

    0/2

    The first agentic lessons. What Is Agentic AI? gives the definition everything else rests on — a system is agentic when the model chooses the next action from what it observed, inside a loop with a termination condition — and sorts traditional software, LLM app, RAG and agent by who owns the control flow. Choosing the Right Abstraction turns that into the escalation ladder from plain code to multi-agent that every later stage climbs one rung at a time.

    Before moving on: State the four parts that make a system agentic, tell a workflow from an agent however it is marketed, and place a problem on the ladder from plain code to multi-agent before writing any of it.

  5. 5

    Prompt & Context Engineering

    0/5

    The model only sees what you put in the prompt, so constructing that set of information is the core skill. Selection and compression decide what goes in, ordering matters because of Context Ordering & Lost in the Middle, and Token Budgets keep the whole thing inside the window. It comes before tools and RAG because both of them are ways of filling the context.

    Before moving on: Assemble the exact set of information a model sees on each call, stay inside a token budget, and explain why a fact buried in the middle of a long prompt is often ignored.

    Needs first:LLM Fundamentals
  6. 6

    Structured Outputs

    0/2

    Before a model can drive code it has to return data the code can trust. Structured Outputs shows how to force schema-conformant JSON, and Argument Validation treats the result as untrusted input and validates it at the boundary. Tool calling is structured output with a function attached, so this stage comes first.

    Before moving on: Force a model to return schema-conformant JSON, validate it at the boundary, and handle a validation failure instead of trusting the shape of the text.

  7. 7

    Tool Calling

    0/7

    The step from "model that answers" to "model that acts". Schemas the model selects correctly, argument validation, errors, timeouts and retries, Idempotency for side effects, Parallel vs Sequential Tool Calls and Tool Permissions and Least Privilege. The APIs stage supplied the network discipline; this stage puts a model in front of it.

    Before moving on: Define tool schemas the model picks correctly, validate arguments, handle errors, timeouts and retries, make side effects idempotent, and grant each tool the minimum permissions it needs.

  8. 8

    RAG

    0/8

    Retrieval is how an agent gets knowledge that does not fit in the prompt. The full pipeline in order: parsing and chunking, Embeddings, vector storage, Dense, Sparse & Hybrid Retrieval, metadata filtering, Reranking, and grounding with citations. RAG Evaluation closes the stage by measuring retrieval quality separately from answer quality, which is the habit the evaluation stage generalises.

    Before moving on: Build the full retrieval pipeline from chunking to reranked, cited answers, and measure retrieval quality separately from answer quality.

  9. 9

    Agent Loops

    0/5

    With tools and retrieval in hand you can close the loop. The Agent Loop implements observe-reason-act without a framework, Single Agent and Agent + RAG are the two shapes it usually takes, and Planning Strategies with When Planning Helps compare direct execution, plan-then-execute and ReAct so you can tell when planning only adds latency and cost.

    Before moving on: Implement the observe-reason-act loop without a framework and choose between direct execution, plan-then-execute and ReAct for a given task.

    Needs first:Tool CallingRAG
  10. 10

    State & Memory

    0/3

    A loop that runs for more than one turn needs somewhere to keep what it learned. Memory Types separates context, short-term and long-term memory, Memory Architectures shows how semantic, episodic and procedural memory are stored and retrieved, and Memory Pitfalls explains why storing more usually makes an agent worse.

    Before moving on: Separate context, short-term and long-term memory, decide what deserves persisting, and explain why storing more makes agents worse rather than better.

  11. 11

    MCP

    0/4

    The Model Context Protocol standardises how tools, resources and prompts are exposed to a model, so this stage follows tool calling directly. MCP Primitives: Tools, Resources, Prompts and MCP Lifecycle: Init, Discovery, Invocation cover the protocol and its connection and auth handling; MCP vs Direct Integration vs Function Calling is the decision you will actually be asked to make.

    Before moving on: Build and consume an MCP server, handle its lifecycle and auth, and argue when a direct integration or plain function calling is the better choice.

    Needs first:Tool Calling
  12. 12

    Workflow Orchestration

    0/4

    Most business processes are better served by an explicit state graph than by an open-ended agent. Workflow State Graph models the steps, checkpoints and routing, Router Architecture dispatches to the right branch, and Architecture Tradeoffs with What Agentic Engineers Build put the loop from the previous stages next to the deterministic alternative.

    Before moving on: Model a multi-step process as an explicit state graph with checkpoints and routing, and say when a deterministic workflow beats an open-ended agent.

    Needs first:Agent Loops
  13. 13

    Multi-Agent Systems

    0/8

    Several loops cooperating: supervisor, pipeline, hierarchical and swarm topologies, and the messages the agents exchange in Agent-to-Agent Communication. It comes after workflows because a multi-agent system is a workflow whose nodes are agents. When Not to Use Multi-Agent is the most important lesson here: a single agent with better tools is usually the right answer.

    Before moving on: Choose between supervisor, pipeline, hierarchical and swarm topologies, design the messages agents exchange, and recognise when a single agent with better tools is enough.

  14. 14

    Evaluation

    0/6

    A probabilistic system cannot be unit-tested in the usual way, so this stage teaches how to test it anyway: Golden Datasets, metrics that match the failure you care about, LLM-as-Judge combined with Deterministic Evaluators, and Regression Gates and Online Evaluation before every prompt or model change. You need a working agent to evaluate, which is why it sits after the loop.

    Before moving on: Build a golden dataset, pick metrics that match the failure you care about, combine deterministic evaluators with LLM judges, and run regression evals before every change.

    Needs first:Agent LoopsRAG
  15. 15

    Observability

    0/3

    When a run goes wrong, the trace is how you find out where. Tracing Agents records every LLM call, tool call and state change as spans, Trace Inspection: Debugging from a Trace is the debugging skill of reading one, and Logging, Metrics and Alerts turns token, cost, latency and error counts into alerts. Everything from here on assumes you can see what the agent did.

    Before moving on: Trace every LLM call, tool call and state change as spans, read a trace to find the step where a run went wrong, and alert on token, cost, latency and error rates.

  16. 16

    Guardrails & Security

    0/7

    An agent with tools and retrieval has an attack surface that a chat model does not. Prompt Injection and Indirect Prompt Injection show how retrieved and tool-returned content becomes an instruction, Tool Misuse and Data Exfiltration and Permissions, Authentication and Authorisation limit the damage, and Input and Output Guardrails with Secrets and Untrusted Output close the remaining gaps. It needs both tool calling and RAG because those are the two doors.

    Before moving on: Treat all retrieved and tool-returned content as untrusted data, scope tool permissions to the task, keep secrets out of the context, and put guardrails on both inputs and outputs.

    Needs first:Tool CallingRAG
  17. 17

    Human-in-the-Loop

    0/4

    Some actions should not run without a person. Approval Gates and Risk Classes classifies actions by reversibility and inserts gates where it is low, In-the-Loop vs On-the-Loop and Escalation separates approving from monitoring, and Confidence Thresholds route only the cases that matter to a human. Gates are checkpoints in a workflow graph, and the risk model comes from the security stage.

    Before moving on: Classify actions by risk, insert approval gates where reversibility is low, and design escalation with confidence thresholds so humans review the cases that matter instead of every case.

  18. 18

    Reliability

    0/4

    Agent runs fail in their own ways: infinite loops, context overflow, provider outages, arguments that drift with each retry. Failure Scenarios: Detection and Runbooks names them, Budgets, Limits and Termination caps a run before it does damage, and Fallbacks, Caching and Model Routing keeps the system up when a model or provider is not. Tracing comes first because you cannot fix a failure you cannot see.

    Before moving on: Enumerate the ways an agent run fails and name a mitigation for each: budgets, termination criteria, fallbacks, caching and routing.

  19. 19

    Production Agent Systems

    0/5

    The capstone: everything above applied to a real business process. It revisits Fallbacks, Caching and Model Routing and Logging, Metrics and Alerts in the context of a system that has to stay up and stay cheap, Business Process Automation shows what real deployments look like, and Frameworks: What They Abstract and What It Costs ends with a defensible choice of framework, or none at all.

    Before moving on: Take an agent from demo to production with fallback chains, cost controls, metrics and alerts, and defend the choice of framework or of no framework.