Lab 03
Physical twin of the lab graph
Teaching map only. This playground still calls Ollama on localhost — no AWS keys, no hosted models, no prompts leave the machine.
Browser to edge, then named lab steps inside a region, then metrics and egress.
Browser
Edge
Region
LangGraph physical architecture on AWS
Step Functions
Named states, same order as the lab
Nova Micro
AZ-a / AZ-b
metrics only
TTL session
Browser
egress
The same twin, grouped the way the live Flow column lights.
The front door. Users never talk to a model directly. Traffic hits a CDN and firewall, then a managed HTTP API that starts the state machine — the production twin of this LangGraph.
Decide what kind of question this is before spending a large model. Cheap rules first, then a small classifier if the words are ambiguous.
Turn the question into a vector, then fetch nearby passages. On AWS the corpus lives in S3; Bedrock Knowledge Bases owns chunking, embedding, and query.
Write a first answer from the retrieved passages only. This is the expensive token step, so you run it in more than one Availability Zone and pick a mid-size model.
A second pass that must not invent facts. Guardrails check groundedness; a stronger model writes the critique; Lambda can enforce hard rules (citations present, length).
Apply the critique, write the user-facing answer, then filter it again on the way out. This is the last model call before API Gateway streams the result.
Hand the answer back and keep only operational numbers. Sessions can live in a TTL cache the way this playground keeps state in process memory.