๐Ÿค– Lab 3: Design Your Agent Workflow

Pick a real workflow from your team, choose a workflow pattern, and design the full agent using the Agent Design Canvas.

โฑ 45 minutes + bonus

Exercise Overview

This is the capstone exercise of the workshop. You've learned what GenAI can do, how to communicate with it effectively, and how to build reusable skills and automation. Now you'll put it all together โ€” designing an AI agent for a real workflow from your own team.

You'll work in teams of 3-4 to select a workflow, choose an automation pattern, and fill out an Agent Design Canvas โ€” a structured one-page document that captures everything needed to build the agent: its role, workflow steps, data requirements, guardrails, escalation rules, and expected business impact. This is the document you'd hand to your tech team to start implementation.

๐ŸŽฏ Deliverable

Each team produces an Agent Design Canvas โ€” saved as a markdown file in your workspace. Submit it for LLM-as-Judge scoring โ€” an automated evaluation technique applied to your agent design.
StepWhat you doDuration
Step 1Pick your workflow3 min
Step 2Confirm the workflow pattern2 min
Step 3Design the agent (canvas auto-fills โ†’ paste into Kiro)25 min
Step 4Generate a presentable HTML canvas5 min
Step 5Submit for LLM-as-Judge scoring5 min
โญ BonusBuild it yourself โ€” turn your canvas into Skills + Hooks15 min

Step 1: Pick Your Workflow

Click a workflow below to select it (or pick your own). Your selection will auto-fill the canvas prompt in Step 3.

๐Ÿฅ Claims triage & adjudication
Classify claim type โ†’ validate coverage โ†’ auto-decide or escalate
Recommended: Routing
๐Ÿ“‹ Underwriting risk assessment
Assess medical + financial + compliance in parallel โ†’ risk classification
Recommended: Parallelization
๐Ÿ“œ Regulatory change impact
Scan MAS circular โ†’ assess gaps โ†’ generate action items
Recommended: Chaining
๐Ÿ” Audit finding management
Track findings โ†’ assess remediation โ†’ escalate overdue โ†’ board report
Recommended: Orchestration
โœ… QA test case generation
Read requirements โ†’ generate test cases โ†’ prioritize โ†’ coverage report
Recommended: Chaining
๐Ÿ“Š Monthly executive reporting
Pull KPIs โ†’ analyze trends โ†’ generate narrative โ†’ create deck
Recommended: Chaining
๐Ÿ›ก๏ธ Security incident triage
Detect alert โ†’ classify severity โ†’ gather context โ†’ route response
Recommended: Routing
๐Ÿ”„ Policy renewal review
Identify renewals โ†’ assess coverage gaps โ†’ recommend to advisor
Recommended: Chaining
โš ๏ธ Best choice: Pick a workflow that your team does at least weekly and that takes at least 30 minutes each time. The bigger the time savings, the stronger your business case. You can also pick your own workflow โ€” just type it into the canvas prompt in Step 3.

Step 2: Confirm the Pattern

The recommended pattern is pre-selected based on your workflow choice. You can change it if you prefer a different approach:

๐Ÿ”— Chaining
A โ†’ B โ†’ C (sequential)
โšก Parallelization
Run 3 analyses โ†’ combine
๐Ÿ”€ Routing
Classify โ†’ route to correct path
๐ŸŽฏ Orchestration
Decision points + human review

Step 3: Fill the Agent Design Canvas

The canvas below is pre-filled based on your selections. Review it, adjust anything that doesn't fit your team's specifics, then copy and paste into Kiro:

AGENT DESIGN CANVAS โ€” Copy & paste into Kiro
Help me design an AI agent using this canvas template. I'll describe my workflow and you help me fill in each section. # Agent Design Canvas ## Agent Identity - **Name:** [descriptive name for the agent] - **Role:** [Use the persona formula: Title + Experience + Specialty + Characteristic + Behavior] - **Trigger:** [What event starts this agent? e.g., "New file uploaded", "Weekly schedule", "Manual request"] ## Workflow - **Pattern:** [ ] Chaining [ ] Parallelization [ ] Routing [ ] Orchestration - **Steps:** 1. [First action โ€” what skill runs? what does it produce?] 2. [Second action โ€” what does it do with the output of step 1?] 3. [Third action โ€” decision point, human review, or final output?] ## Data - **Inputs:** [What data does the agent need? Files, databases, APIs?] - **Outputs:** [What does it produce? Report, notification, decision?] - **Knowledge base:** [What documents should it reference? Policies, guidelines?] ## Guardrails - **Must NOT:** [List negative constraints โ€” what the agent must never do] - **Escalate when:** [When does a human take over? Thresholds, confidence levels?] - **Data rules:** [Grounding rules โ€” only use provided data? Citation requirements?] ## Kiro Implementation - **Steering file:** [What global rules apply?] - **Skills needed:** [List each SKILL.md file needed for the workflow steps] - **Hook trigger:** [What event triggers the workflow?] - **MCP connections needed:** [What external systems? e.g., database, Slack, email, Google Drive โ€” your tech team sets these up] ## Business Impact - **Current process:** [How is this done today? How long does it take?] - **With agent:** [Expected time savings per week/month] - **Success metric:** [How will you measure if the agent is working?] --- My workflow is: [Select a workflow in Step 1 above] The pattern I chose is: [Select a pattern in Step 2 above] My team currently spends [TIME] on this each [FREQUENCY]. Save the completed canvas as a markdown file at lab3-agent-canvas/agent-design-canvas.md
๐Ÿ’ก Tips for a strong canvas:
  • Be specific about the trigger โ€” "new CSV file in /data" is better than "when needed"
  • Name each skill โ€” each step should map to a SKILL.md file
  • Include escalation rules โ€” when does a human take over?
  • Quantify the business impact โ€” "saves 10 hours/week" is more compelling than "saves time"
  • Use proven techniques โ€” persona for the role, structured output for each skill, negative constraints for guardrails

Step 4: Generate a Presentable Canvas

Turn your Agent Design Canvas into a polished, visual HTML page that you can present on screen. In the same session, paste:

PROMPT โ€” Generate HTML Canvas
Based on the Agent Design Canvas we just created, generate a single-page HTML file called "lab3-agent-canvas/agent-canvas.html" in the workspace. The page should be a professional, presentation-ready visual layout of the canvas with: 1. A header bar with the agent name, workflow pattern badge, and team name 2. A visual workflow diagram showing the steps as connected boxes with arrows (use CSS โ€” no images needed). Color-code by step type: blue for data input, green for AI processing, orange for decision points, red for human escalation 3. A two-column layout below the diagram: - Left column: Agent Identity (name, role, trigger), Data (inputs, outputs, knowledge base) - Right column: Guardrails (must-not rules, escalation triggers, data rules), Business Impact (current vs with agent, success metrics) 4. A "Kiro Implementation" section at the bottom showing: steering file, skill files needed, hook trigger, MCP connections โ€” styled as a technical spec card with a dark background 5. A footer: "Agent Design Canvas โ€” AnyCompany Insurance โ€” [Team Name]" Design: Clean white background, AnyCompany red (#E31837) accent, professional typography, subtle shadows. Should look good projected on a screen. Open the HTML file in the browser after generating.
โœ… Checkpoint: You should now have:
  • A completed Agent Design Canvas (markdown โ€” lab3-agent-canvas/agent-design-canvas.md)
  • A visual HTML version (lab3-agent-canvas/agent-canvas.html) ready to present on screen
๐Ÿ’ก From canvas to production โ€” see it realized in Lab 4

Want to see what happens when an Agent Design Canvas becomes real working code? Lab 4: Agentic Claims Processing โ†’ takes the Claims Processing canvas and implements it as a multi-agent system using the Strands Agents SDK:
Canvas SectionWhat it becomes in code
Agent Identity โ†’ RoleSystem prompt for each agent (system_prompt="You are a Senior Medical Claims Reviewer...")
Workflow โ†’ StepsThe pipeline: OBSERVE โ†’ PLAN โ†’ ACT โ†’ REFLECT
Workflow โ†’ PatternTriage Agent (Routing) + ThreadPoolExecutor (Parallelization)
Data โ†’ Inputs@tool functions: read_claim(), check_policy()
Guardrails โ†’ Must NOTFraud indicators in check_fraud_indicators() tool
Guardrails โ†’ Escalate whenThreshold checks: if policy_age < 12: "FLAG"
Kiro Implementation โ†’ Skills4 Strands Agents: Triage, Medical, Fraud, Adjudicator

The canvas is the design document. Lab 4 is the working prototype. This is exactly the handoff your tech team would do โ€” your canvas tells them what to build, the code shows how it works.

Step 5: Submit & Score Your Canvas

Submit your Agent Design Canvas for automated LLM-as-Judge scoring. Paste the markdown content from your canvas file, and the AI will evaluate it on 6 criteria.

๐Ÿ“ค Submit Your Canvas

Submit as a team or individually. Resubmitting with the same name replaces your previous entry.

๐Ÿ’ก How scoring works

Your canvas is sent to Amazon Bedrock (Claude Sonnet) which acts as an LLM-as-Judge โ€” an automated evaluation technique. Nothing is hidden: the judge is given a fixed rubric of 6 criteria, scores each one from 1 to 5, and adds them up for a total out of 30. It runs at a low temperature (0.1) so scoring stays consistent between runs.

What makes a strong canvas? The same principles from the workshop apply โ€” be specific (not vague), ground everything in data (not assumptions), include guardrails (not just the happy path), and quantify the business impact (not just "saves time"). The more concrete and actionable your canvas, the higher it scores.

๐Ÿ” Exactly what the judge evaluates

These are the 6 criteria โ€” verbatim โ€” that the AI judge is instructed to score. Full transparency: this is the same rubric running behind the leaderboard.

CriterionPointsWhat the judge looks for
1. Problem Clarity1โ€“5Is the workflow clearly defined with measurable pain points?
2. Pattern Fit1โ€“5Does the chosen pattern (chaining, routing, parallel, etc.) match the workflow? Are the steps logical?
3. Guardrails & Safety1โ€“5Are negative constraints specific? Are escalation thresholds quantified?
4. Implementation Readiness1โ€“5Could a tech team build from this? Are skills, hooks, and steering specified?
5. Business Impact1โ€“5Is ROI quantified with measurable success metrics?
6. AnyCompany Relevance1โ€“5Grounded in AnyCompany Insurance's domain โ€” life/health/motor/travel, claims, underwriting, policy servicing, MAS regulations, the Singapore market?

The judge is deliberately strict โ€” a 5 means genuinely exceptional. Alongside the scores you'll get written strengths and improvements so you know exactly why you landed where you did.

๐Ÿ How your total maps to a verdict

ScoreVerdictWhat it means
25-30๐Ÿฅ‡ Production-ReadyYour tech team could start building from this canvas this week
19-24๐Ÿฅˆ Strong DesignSolid foundation โ€” address the feedback and it's ready
13-18๐Ÿฅ‰ Good StartRight direction but needs more specificity โ€” check the improvements
Below 13๐Ÿ”„ Needs ReworkAdd more detail โ€” the AI feedback tells you exactly where

๐Ÿ† Leaderboard

All team submissions ranked by score. Click "Refresh" to see new entries.

Loading leaderboard...

โญ Bonus: Turn Your Canvas into a Working Agent

You've designed the canvas and submitted it for scoring. Now let's close the loop โ€” see how a canvas becomes a working Kiro implementation using steering, skills, and hooks. No Python code needed โ€” everything runs through Kiro's native features.

๐ŸŽฏ The demo

We'll use the Underwriting Risk Assessment canvas as the example. The package includes 3 synthetic applications (low/medium/high risk), underwriting guidelines, 4 specialist skills, a steering file, and a pipeline hook. You install it, click the hook, and watch Kiro assess applications like a team of underwriters.

How the Canvas Becomes Real Files

Remember the Agent Design Canvas you filled out in Steps 3-4? Here's exactly how each section of that canvas translates into Kiro implementation files. This is the same process your tech team would follow:

Canvas Section Becomes This File What It Contains
Guardrails โ†’ Must NOT
Guardrails โ†’ Escalate when
Guardrails โ†’ Data rules
.kiro/steering/underwriting-rules.md Authority limits (no auto-approve > $2M), escalation triggers (age > 60, PEP, cardiac event), data grounding rules ("cite specific data points")
Workflow โ†’ Step 1
Agent Identity โ†’ Role
.kiro/skills/medical-assessment/SKILL.md Persona: "Senior Medical Underwriter, 15 years experience" โ€” extracts BMI, pre-existing conditions, family history โ†’ outputs risk score
Workflow โ†’ Step 2 .kiro/skills/financial-assessment/SKILL.md Checks income-to-coverage ratio, over-insurance, source of funds โ€” uses guideline multiples (10x for <50, 8x for 50-60)
Workflow โ†’ Step 3 .kiro/skills/compliance-check/SKILL.md Sanctions screening, PEP status, AML check โ€” "Sanctions match = AUTOMATIC BLOCK (no exceptions)"
Workflow โ†’ Step 4 (Synthesize) .kiro/skills/underwriting-decision/SKILL.md Chief Underwriter persona โ€” combines all 3 assessments, applies "most conservative" rule, outputs final APPROVE/DECLINE/REFER
Kiro Implementation โ†’ Hook trigger
Workflow โ†’ Pattern
.kiro/hooks/underwriting-pipeline.json One-click trigger that orchestrates: read guidelines โ†’ medical โ†’ financial โ†’ compliance โ†’ decision synthesis โ†’ save report
Canvas โ†’ Implementation Flow
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚
  ๐Ÿ“ Agent Design Canvas (your markdown from Step 3-4)
โ”‚
                        โ–ผ
  ๐Ÿ›ก๏ธ Guardrails section โ†’ steering/underwriting-rules.md
                        โ–ผ
  โš™๏ธ Workflow steps โ†’ skills/ (one SKILL.md per step)
      โ”œโ”€โ”€ Step 1 โ†’ medical-assessment/SKILL.md
      โ”œโ”€โ”€ Step 2 โ†’ financial-assessment/SKILL.md
      โ”œโ”€โ”€ Step 3 โ†’ compliance-check/SKILL.md
      โ””โ”€โ”€ Step 4 โ†’ underwriting-decision/SKILL.md
                        โ–ผ
  ๐Ÿ”˜ Hook trigger โ†’ hooks/underwriting-pipeline.json
                        โ–ผ
  โ–ถ๏ธ Click hook โ†’ Agent runs full pipeline
โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
๐Ÿ’ก Key insight

The Agent Design Canvas is not just a workshop exercise โ€” it's a blueprint. Every section maps directly to a file your tech team creates. The canvas tells them what to build; the files below show how it looks when built. Let's install them and see it run.

Step 1: Download & Install the Underwriting Package

  1. Download lab3-bonus-underwriting.zip and extract it into your workspace
  2. Copy the Kiro files into place:
    KIRO PROMPT
    Copy the Kiro configuration files from lab3-bonus-underwriting into the right locations: - Copy kiro-files/steering/underwriting-rules.md to .kiro/steering/underwriting-rules.md - Copy kiro-files/skills/medical-assessment/ to .kiro/skills/medical-assessment/ - Copy kiro-files/skills/financial-assessment/ to .kiro/skills/financial-assessment/ - Copy kiro-files/skills/compliance-check/ to .kiro/skills/compliance-check/ - Copy kiro-files/skills/underwriting-decision/ to .kiro/skills/underwriting-decision/ - Copy kiro-files/hooks/underwriting-pipeline.json to .kiro/hooks/underwriting-pipeline.json
โœ… Checkpoint: You should now see in your workspace:
  • .kiro/steering/underwriting-rules.md โ€” guardrails (escalation thresholds, data grounding)
  • .kiro/skills/ โ€” 4 skills (medical, financial, compliance, decision synthesis)
  • .kiro/hooks/underwriting-pipeline.json โ€” the pipeline trigger button
  • lab3-bonus-underwriting/sample-data/ โ€” 3 applications + guidelines

Step 2: Run the Pipeline (Low Risk โ€” APP-2025-001)

Click the "Underwriting Risk Assessment Pipeline" hook button in the Kiro hooks panel, then tell it which application to process:

KIRO PROMPT โ€” After clicking the hook
Process application APP-2025-001 from lab3-bonus-underwriting/sample-data/applications/APP-2025-001.md Read the underwriting guidelines first, then run all four assessments in order.
๐Ÿ‘€ What to watch for:
  • Kiro reads the guidelines first (grounding)
  • Medical: BMI 20.9 (normal), no conditions, non-smoker โ†’ LOW risk
  • Financial: Income multiple 4.8x (within 10x limit) โ†’ LOW risk
  • Compliance: All clear, no PEP, no AML flags โ†’ CLEAR
  • Decision: Should be APPROVE โœ… (Preferred)

Step 3: Run a Risky Application (APP-2025-003)

Now try the high-risk case โ€” this one should trigger multiple escalation rules:

KIRO PROMPT
Process application APP-2025-003 from lab3-bonus-underwriting/sample-data/applications/APP-2025-003.md Read the underwriting guidelines first, then run all four assessments in order.
๐Ÿ”ด Expected: Multiple red flags
  • Medical: Cardiac stent (2023), uncontrolled diabetes (HbA1c 8.2%), obese, smoker โ†’ VERY HIGH
  • Financial: Income multiple 42x (exceeds 25x guideline!) โ†’ HIGH (over-insurance)
  • Compliance: Former PEP + single premium $500K โ†’ Enhanced due diligence required
  • Decision: Should be REFER TO SENIOR UW ๐ŸŸก or DECLINE ๐Ÿ”ด

Step 4: Compare the Decisions

Check the output reports:

KIRO PROMPT
Read both assessment reports in lab3-bonus-underwriting/output/ and create a comparison table showing how the same pipeline produced different decisions for different risk profiles. Highlight which specific data points triggered the escalation for APP-2025-003.
๐ŸŽ‰ What just happened

You saw a Parallelization pattern in action โ€” three specialist assessments (medical, financial, compliance) feeding into a synthesis decision. All powered by:
  • Steering โ€” guardrails that apply to every assessment (escalation thresholds, data grounding)
  • Skills โ€” 4 specialist personas with structured output formats
  • Hook โ€” one-click trigger that orchestrates the full pipeline
  • No code โ€” everything is prompt-based, using Kiro's native features

This is the same architecture your tech team would build โ€” but you just demonstrated it in 5 minutes with zero Python.

๐Ÿ”ฌ Optional: Try APP-2025-002 (Medium Risk)

If time permits, run the medium-risk application. This one is interesting because it's borderline โ€” controlled diabetes, ex-smoker, overweight. The agent should recommend Standard with conditions or Substandard with loading. Watch how the agent weighs competing factors.

KIRO PROMPT
Process application APP-2025-002 from lab3-bonus-underwriting/sample-data/applications/APP-2025-002.md
๐Ÿ”— Want to see the code version?

Lab 4: Agentic Claims Processing โ†’ takes a similar multi-agent design and implements it as a Python system using the Strands Agents SDK. Same patterns (routing + parallelization), different implementation path โ€” Skills + Hooks for prototyping, Strands SDK for production.