Agentic AI Foundation โ AnyCompany Insurance ยท Thank you for a great workshop!
๐ Important Links & Access
S for speaker notes, F for fullscreen.๐ What We Covered
| Module | Topic | Key Takeaway |
|---|---|---|
| M0 | Introduction | Workshop goals, environment setup, what agentic AI means for insurance |
| M1 | From LLMs to Agents | Fallacy of composition โ chatbots aren't agents. Agents need perception, reasoning, and action. |
| M2 | Exploring Agentic AI | 4 agent types: Workflow, Autonomous, Hybrid, Multi-Agent. 4 autonomy levels (L1โL4). |
| M3 | Agentic Workflows โญ | Chaining, Parallelization, Routing, Orchestration โ the 4 patterns that power every agent workflow |
| M4 | AWS Developer Tools | Bedrock Agents, AgentCore, Kiro IDE, MCP integration โ the AWS AI landscape |
| M5 | Frameworks | Strands Agents SDK, LangGraph, CrewAI โ when to use which framework |
| M6 | Custom Solutions | Production patterns, guardrails, evaluation, human-in-the-loop design |
| M7 | Wrap Up | Key takeaways, organizational readiness, next steps for your team |
๐ The 4 Workflow Patterns
โก Key Concepts Reference
๐๏ธ The Kiro Automation Stack
| Layer | What It Does | Who Owns It |
|---|---|---|
| Steering | Global rules that apply to every AI interaction โ compliance standards, formatting, authority limits | Team leads / policy owners |
| Skills | On-demand expertise that activates for specific tasks โ reusable prompt templates with structured outputs | Domain experts |
| Hooks | Automated triggers โ run actions when files change, on schedule, or on manual trigger (one-click pipelines) | Process owners |
| MCP | Data connections โ connect AI to databases, APIs, and external tools via standardized protocol | Tech team |
๐ค Agent Design Canvas โ Main Exercise
The Agent Design Canvas captures everything needed to build an AI agent for your workflow:
| Canvas Section | What It Defines | Who Uses It |
|---|---|---|
| Agent Identity | Name, role (persona), trigger event | Everyone โ sets the scope |
| Workflow | Pattern choice + step-by-step process | Tech team โ architecture |
| Data | Inputs, outputs, knowledge base | Tech team โ integrations |
| Guardrails | Must-NOT rules, escalation triggers, data rules | Risk/Compliance โ governance |
| Kiro Implementation | Steering, skills, hooks, MCP connections | Tech team โ build plan |
| Business Impact | Current process time, expected savings, success metrics | Leadership โ ROI case |
๐ฏ Interactive Explainers (Self-Paced Reference)
These stay available after the workshop. Use them to refresh concepts or share with colleagues who weren't in the room.
๐ Recommended Next Steps
๐ฌ Participant Q&A โ Parking Lot
Questions raised during the workshop, answered with technical accuracy and practical guidance.
Best practice: One skill per persona/task. Each SKILL.md should focus on a single, well-defined responsibility โ same principle as microservices.
| Approach | Pros | Cons |
|---|---|---|
| One skill per persona | Easier to test, clearer activation triggers, smaller token footprint, reusable across workflows | More files to manage |
| Combined mega-skill | Fewer files | Harder to debug, larger token usage on every activation, ambiguous triggers |
Why: Kiro skills load full content on-demand when the agent determines relevance. A focused skill (500-1000 words) loads fast and keeps context clean. A mega-skill (5000+ words) wastes tokens when only one capability is needed.
claims-triage/SKILL.md (router) + medical-claims-review/SKILL.md (specialist) + fraud-investigation/SKILL.md (specialist) โ each independently testable, each with a clear activation trigger.Yes โ all three have production-ready MCP servers available today.
| Tool | MCP Server | Auth Method |
|---|---|---|
| Jira + Confluence | Official Atlassian MCP Server (GA) | OAuth 2.0 |
| Jira + Confluence | mcp-atlassian (community) | API token or OAuth |
| Figma | Official Figma MCP Server | OAuth 2.0 |
| Figma | Figma-Context-MCP (community) | Personal access token |
What you can do: Jira โ create/update issues, search with JQL, transition tickets. Confluence โ read/create/update pages, search content. Figma โ read design files, extract component properties, get layout info.
Not directly โ hooks do not chain to each other by design. This prevents infinite loops and makes debugging straightforward.
Indirect chaining is possible:
| Method | How It Works |
|---|---|
| Agent-mediated | Hook A uses askAgent โ agent creates a file โ file creation triggers Hook B (fileCreated) |
| Single orchestrator | One hook with a comprehensive prompt that runs multiple skills in sequence |
| Command chaining | Hook uses runCommand with a script that performs multiple steps |
userTriggered hook with an askAgent prompt that orchestrates the full pipeline. This is what the underwriting bonus exercise does โ one hook triggers the agent to run 4 skills in sequence.Best practice: AI automates dev/staging; humans approve production deployments. This follows the principle of progressive trust.
| Environment | AI Agent Role | Human Role |
|---|---|---|
| Dev/Sandbox | Full autonomy โ write code, run tests, auto-deploy, fix lint errors | Review when needed |
| Staging | Assisted โ run tests, security scans, performance tests, auto-deploy to staging | Review flagged issues |
| Production | Read-only โ generate change request, provide risk assessment, recommend rollback | Approve and deploy |
Where AI adds value without deploying to prod: Code review & security scanning, test generation, IaC validation, dependency auditing, change impact analysis, post-deploy anomaly detection.
Bedrock = managed agent platform. Direct = raw model API access.
| Aspect | Amazon Bedrock | Direct Provider |
|---|---|---|
| What you get | Agent orchestration, tool use, knowledge bases, guardrails, memory | Raw inference (text in โ text out) |
| Security | Data stays in your AWS account, VPC endpoints, IAM, CloudTrail | Data sent to provider's endpoint |
| Guardrails | Native โ content filtering, PII redaction, topic denial, grounding | You build your own |
| Model choice | 30+ models via single Converse API (swap via config) | Only that provider's models |
| Compliance | SOC 2, ISO 27001, HIPAA, PCI DSS, MAS TRM compatible | Varies by provider |
| Vendor lock-in | Low โ same API for all models | Locked to provider's API format |
Yes โ Kiro employs several context management strategies:
| Technique | How It Works |
|---|---|
| Conversation compaction | When context window approaches its limit, earlier exchanges are summarized while preserving key decisions and recent context |
| On-demand skill loading | Only YAML frontmatter (name + description) loads at startup (~50-100 tokens). Full content loads only when relevant. |
| Conditional steering | Steering files with inclusion: fileMatch or inclusion: manual only load when relevant, not always |
| Intelligent file reading | Large files use AST parsing to extract only relevant signatures rather than loading entire files |
| Search-based context | Targeted search (grep, file search) finds only relevant sections rather than loading entire codebases |
What Kiro does NOT do: It does not compress text in a lossy way or use a separate summarizer model to pre-process prompts. The actual tokens sent to the model are standard text.
inclusion: manual for large reference docs, and break large steering files into conditionally-included pieces.Hallucination risk increases with conversation length due to context dilution โ earlier facts get de-prioritized as the context window fills. Proven mitigation strategies:
| Strategy | How It Helps |
|---|---|
| RAG (Retrieval-Augmented Generation) | Ground every response in retrieved source documents. Re-retrieve for each new question rather than relying on earlier context. |
| Structured output enforcement | Require citations: "Based on [document, section X]: ..." Forces grounding over generation. |
| Data grounding rules | In steering: "Only use provided documents. State INFORMATION NOT PROVIDED for missing data." |
| Low temperature (0.0-0.2) | Reduces creative generation for factual tasks. Higher temperature only for brainstorming. |
| Chunked verification (Chaining) | Break tasks into steps, verify each output before proceeding. Each step produces verifiable results. |
| Multi-agent consensus | Multiple agents assess same data independently. Disagreements flag potential hallucinations. |
| Context window management | Start new sessions for new topics. Re-inject critical reference docs periodically. |
| Conversation compaction | Kiro auto-compacts, preserving key facts while summarizing verbose exchanges. |
Agentic QA uses autonomous AI agents that plan, execute, and adapt testing workflows based on goals rather than fixed scripts.
4 maturity levels:
| Level | Approach | What AI Does | Human Role |
|---|---|---|---|
| L1 | AI-Assisted | Generate test cases, suggest assertions, auto-complete scripts | Reviews everything |
| L2 | AI-Augmented | Auto-generate regression suites, self-heal broken selectors, generate test data | Focuses on exploratory testing |
| L3 | Agentic QA | Decides what to test based on code changes, risk, and coverage gaps | Sets goals, reviews results |
| L4 | Self-Healing | Detects failures, diagnoses root cause, fixes tests, re-runs, reports | Handles only novel failures |
Agentic QA workflow patterns:
| Pattern | QA Application | Value |
|---|---|---|
| Chaining | Requirements โ Generate test cases โ Prioritize by risk โ Generate test data โ Output suite | 6 hours manual โ 30 min review |
| Routing | Code change โ Classify (UI/API/DB) โ Route to targeted test suite โ Run โ Report | 4-hour full regression โ 20-min targeted |
| Orchestration | Agent explores app โ Detects anomalies โ Generates bug reports โ Prioritizes โ Escalates | Finds edge cases humans miss; runs 24/7 |
| Parallelization | Failed tests โ Check UI change? + Check API change? + Real bug? โ Auto-fix or file bug | 40% reduction in test maintenance |
๐ Workshop Complete