Lab 02
Physical twin of the lab path
Teaching map only. Inference stays on Ollama at localhost — no cloud keys, no hosted models, no prompts leave the machine.
Browser to edge, then named lab steps inside a region, then metrics and egress.
Browser
Edge
Region
RAG physical architecture on AWS
Step Functions
Named states, same order as the lab
Titan embed
AZ-a / AZ-b
metrics only
Browser
egress
The same twin, grouped the way the live Flow column lights.
Ask lands on a managed API, then a short workflow: embed, retrieve, generate. The twin of POST /labs/rag/ask.
The question is embedded with the same model that built the Chroma index (Ollama nomic-embed-text here). On AWS that is Titan on Bedrock.
Top-k passages from the corpus index. Locally ephemeral Chroma over data/corpus/. On AWS: Bedrock Knowledge Bases on OpenSearch Serverless.
The model is only allowed the retrieved passages. This lab’s prompt says: if they are not enough, say you do not know. That bind is the whole point of RAG.
One completion from the bound passages. Locally llama3.2 via Ollama. On AWS: Bedrock Claude or Llama, on-demand or provisioned.
Citations ([filename]) and a refuse-if-unsupported rule. Guardrails can score groundedness. The UI shows the exact chunks the model was given.
Return the answer plus citation cards. Metrics only — not the question or passages in logs.