๐Ÿ“จ Workshop Summary

Agentic AI Foundation โ€” AnyCompany Insurance ยท Thank you for a great workshop!

ModuleTopicKey Takeaway
M0IntroductionWorkshop goals, environment setup, what agentic AI means for insurance
M1From LLMs to AgentsFallacy of composition โ€” chatbots aren't agents. Agents need perception, reasoning, and action.
M2Exploring Agentic AI4 agent types: Workflow, Autonomous, Hybrid, Multi-Agent. 4 autonomy levels (L1โ†’L4).
M3Agentic Workflows โญChaining, Parallelization, Routing, Orchestration โ€” the 4 patterns that power every agent workflow
M4AWS Developer ToolsBedrock Agents, AgentCore, Kiro IDE, MCP integration โ€” the AWS AI landscape
M5FrameworksStrands Agents SDK, LangGraph, CrewAI โ€” when to use which framework
M6Custom SolutionsProduction patterns, guardrails, evaluation, human-in-the-loop design
M7Wrap UpKey takeaways, organizational readiness, next steps for your team
1. Prompt Chaining
Sequential steps where each output feeds the next. Trades speed for accuracy โ€” each step is testable and debuggable independently.
AnyCompany example: Claim received โ†’ Extract data โ†’ Validate against policy โ†’ Check authority limits โ†’ Approve/Escalate โ†’ Generate decision letter
2. Parallelization
Multiple tasks run simultaneously, then results are aggregated. Faster than sequential and provides multiple perspectives.
AnyCompany example: Underwriting โ€” Medical Assessor + Financial Assessor + Compliance Officer assess simultaneously โ†’ Unified risk classification
3. Routing
A classifier directs input to the right processing pipeline based on content type. Each path is optimized for its category.
AnyCompany example: New claim โ†’ Classify (health/motor/life/travel) โ†’ Route to appropriate specialist workflow
4. Orchestration
The most complex pattern โ€” dynamically manages workflows, spawns subtasks based on conditions, and includes human-in-the-loop at configurable thresholds.
AnyCompany example: Claims adjudication โ€” <$5K auto-approve, $5K-$50K AI+adjuster review, >$50K mandatory senior review. Spawns fraud investigation if indicators detected.
Key insight: Most real workflows combine 2-3 patterns. A life insurance underwriting application uses routing (classify risk tier) + parallelization (simultaneous medical/financial/compliance checks) + chaining (sequential decision synthesis). The Agent Design Canvas helps you map which patterns fit your workflow.
Agentic Loop
Observe โ†’ Plan โ†’ Act โ†’ Reflect โ€” continuous cycle, not linear
Perception โ†’ Reasoning โ†’ Action
The 3 components that make an AI system "agentic"
Fallacy of Composition
What works for parts won't work for the whole โ€” why chatbots aren't agents
Autonomy Levels (L1-L4)
L1: Copilot โ†’ L2: AI-enhanced โ†’ L3: Autonomous โ†’ L4: Fully autonomous
MCP (Model Context Protocol)
Standardized way to connect AI to databases, APIs, and tools
Human-in-the-Loop
Configurable thresholds where humans review AI decisions
Guardrails
Safety boundaries โ€” what the agent must NOT do, when to escalate
Agent Design Canvas
Structured one-page document to hand to your tech team for implementation
LayerWhat It DoesWho Owns It
SteeringGlobal rules that apply to every AI interaction โ€” compliance standards, formatting, authority limitsTeam leads / policy owners
SkillsOn-demand expertise that activates for specific tasks โ€” reusable prompt templates with structured outputsDomain experts
HooksAutomated triggers โ€” run actions when files change, on schedule, or on manual trigger (one-click pipelines)Process owners
MCPData connections โ€” connect AI to databases, APIs, and external tools via standardized protocolTech team
The key distinction: Skills = your domain knowledge (instructions you write). MCP = technical connections (your tech team sets up). You own the "what" and "how to assess." They own the "where the data lives."

What You Designed

The Agent Design Canvas captures everything needed to build an AI agent for your workflow:

Canvas SectionWhat It DefinesWho Uses It
Agent IdentityName, role (persona), trigger eventEveryone โ€” sets the scope
WorkflowPattern choice + step-by-step processTech team โ€” architecture
DataInputs, outputs, knowledge baseTech team โ€” integrations
GuardrailsMust-NOT rules, escalation triggers, data rulesRisk/Compliance โ€” governance
Kiro ImplementationSteering, skills, hooks, MCP connectionsTech team โ€” build plan
Business ImpactCurrent process time, expected savings, success metricsLeadership โ€” ROI case
What to do with your canvas: Share it with your tech team. It's the specification document they need to start building. The more specific your guardrails and escalation rules, the faster they can implement safely.

These stay available after the workshop. Use them to refresh concepts or share with colleagues who weren't in the room.

What to Do This Week

  1. Share your Agent Design Canvas with your tech team or manager. It's the starting point for any AI agent initiative โ€” the "what to build" document.
  2. Identify one workflow in your team that takes >30 minutes and happens weekly. That's your first agent candidate.
  3. Review the use cases page (agentic-usecases.html) โ€” find the use case closest to your team's work and use it as a reference when pitching internally.
  4. Explore the interactive explainers โ€” share the workflow patterns explainer with colleagues who weren't in the workshop.

What to Do This Month

  1. Pilot one agent โ€” start with a low-risk, high-frequency workflow (e.g., report generation, data extraction, document summarization)
  2. Define guardrails first โ€” before building, agree on what the agent must NOT do and when it must escalate to a human
  3. Measure the baseline โ€” document how long the current process takes so you can quantify the improvement
  4. Connect with your Cloud COE โ€” they can help set up Bedrock access, MCP connections, and the technical infrastructure

Questions raised during the workshop, answered with technical accuracy and practical guidance.

Q1 Should each SKILL.md focus on a single persona, or can we combine multiple personas into one skill?

Best practice: One skill per persona/task. Each SKILL.md should focus on a single, well-defined responsibility โ€” same principle as microservices.

ApproachProsCons
One skill per personaEasier to test, clearer activation triggers, smaller token footprint, reusable across workflowsMore files to manage
Combined mega-skillFewer filesHarder to debug, larger token usage on every activation, ambiguous triggers

Why: Kiro skills load full content on-demand when the agent determines relevance. A focused skill (500-1000 words) loads fast and keeps context clean. A mega-skill (5000+ words) wastes tokens when only one capability is needed.

Example structure: claims-triage/SKILL.md (router) + medical-claims-review/SKILL.md (specialist) + fraud-investigation/SKILL.md (specialist) โ€” each independently testable, each with a clear activation trigger.

Q2 Can I create MCP connections to Jira, Confluence, and Figma?

Yes โ€” all three have production-ready MCP servers available today.

ToolMCP ServerAuth Method
Jira + ConfluenceOfficial Atlassian MCP Server (GA)OAuth 2.0
Jira + Confluencemcp-atlassian (community)API token or OAuth
FigmaOfficial Figma MCP ServerOAuth 2.0
FigmaFigma-Context-MCP (community)Personal access token

What you can do: Jira โ€” create/update issues, search with JQL, transition tickets. Confluence โ€” read/create/update pages, search content. Figma โ€” read design files, extract component properties, get layout info.

Security note: Use OAuth 2.0 (not personal API tokens) for production. OAuth provides scoped permissions, token expiry, and audit trails โ€” important for MAS compliance.

Q3 Can a Kiro hook trigger another hook?

Not directly โ€” hooks do not chain to each other by design. This prevents infinite loops and makes debugging straightforward.

Indirect chaining is possible:

MethodHow It Works
Agent-mediatedHook A uses askAgent โ†’ agent creates a file โ†’ file creation triggers Hook B (fileCreated)
Single orchestratorOne hook with a comprehensive prompt that runs multiple skills in sequence
Command chainingHook uses runCommand with a script that performs multiple steps
Recommended: Use a single userTriggered hook with an askAgent prompt that orchestrates the full pipeline. This is what the underwriting bonus exercise does โ€” one hook triggers the agent to run 4 skills in sequence.

Q4 How should AI agents be positioned in CI/CD pipelines from a security perspective? Should AI deploy to production?

Best practice: AI automates dev/staging; humans approve production deployments. This follows the principle of progressive trust.

EnvironmentAI Agent RoleHuman Role
Dev/SandboxFull autonomy โ€” write code, run tests, auto-deploy, fix lint errorsReview when needed
StagingAssisted โ€” run tests, security scans, performance tests, auto-deploy to stagingReview flagged issues
ProductionRead-only โ€” generate change request, provide risk assessment, recommend rollbackApprove and deploy

Where AI adds value without deploying to prod: Code review & security scanning, test generation, IaC validation, dependency auditing, change impact analysis, post-deploy anomaly detection.

For AnyCompany (MAS-regulated): MAS Technology Risk Management (TRM) guidelines require human accountability for production changes. AI can prepare, validate, and recommend โ€” but a named human must approve production deployments. This is a regulatory requirement for financial institutions in Singapore.

Q5 What is the difference between invoking an agent via Amazon Bedrock versus calling a model provider (Anthropic/DeepSeek) directly?

Bedrock = managed agent platform. Direct = raw model API access.

AspectAmazon BedrockDirect Provider
What you getAgent orchestration, tool use, knowledge bases, guardrails, memoryRaw inference (text in โ†’ text out)
SecurityData stays in your AWS account, VPC endpoints, IAM, CloudTrailData sent to provider's endpoint
GuardrailsNative โ€” content filtering, PII redaction, topic denial, groundingYou build your own
Model choice30+ models via single Converse API (swap via config)Only that provider's models
ComplianceSOC 2, ISO 27001, HIPAA, PCI DSS, MAS TRM compatibleVaries by provider
Vendor lock-inLow โ€” same API for all modelsLocked to provider's API format
For AnyCompany: Bedrock is recommended. MAS TRM requires data residency controls, audit trails, and access management โ€” all native to Bedrock. Direct provider access requires building these controls independently.

Q6 Does Kiro use compression or optimization techniques to reduce token consumption?

Yes โ€” Kiro employs several context management strategies:

TechniqueHow It Works
Conversation compactionWhen context window approaches its limit, earlier exchanges are summarized while preserving key decisions and recent context
On-demand skill loadingOnly YAML frontmatter (name + description) loads at startup (~50-100 tokens). Full content loads only when relevant.
Conditional steeringSteering files with inclusion: fileMatch or inclusion: manual only load when relevant, not always
Intelligent file readingLarge files use AST parsing to extract only relevant signatures rather than loading entire files
Search-based contextTargeted search (grep, file search) finds only relevant sections rather than loading entire codebases

What Kiro does NOT do: It does not compress text in a lossy way or use a separate summarizer model to pre-process prompts. The actual tokens sent to the model are standard text.

Practical tip: Write focused skills (fewer tokens per activation), use inclusion: manual for large reference docs, and break large steering files into conditionally-included pieces.

Q7 How do you prevent model hallucination during long conversation sessions?

Hallucination risk increases with conversation length due to context dilution โ€” earlier facts get de-prioritized as the context window fills. Proven mitigation strategies:

StrategyHow It Helps
RAG (Retrieval-Augmented Generation)Ground every response in retrieved source documents. Re-retrieve for each new question rather than relying on earlier context.
Structured output enforcementRequire citations: "Based on [document, section X]: ..." Forces grounding over generation.
Data grounding rulesIn steering: "Only use provided documents. State INFORMATION NOT PROVIDED for missing data."
Low temperature (0.0-0.2)Reduces creative generation for factual tasks. Higher temperature only for brainstorming.
Chunked verification (Chaining)Break tasks into steps, verify each output before proceeding. Each step produces verifiable results.
Multi-agent consensusMultiple agents assess same data independently. Disagreements flag potential hallucinations.
Context window managementStart new sessions for new topics. Re-inject critical reference docs periodically.
Conversation compactionKiro auto-compacts, preserving key facts while summarizing verbose exchanges.
For AnyCompany: Combine RAG (grounding in policy documents) + structured output (cite section numbers) + escalation rules (flag low-confidence answers). This provides the strongest protection in regulated workflows.

Q8 What is the agentic AI strategy for QA use cases from a testing perspective?

Agentic QA uses autonomous AI agents that plan, execute, and adapt testing workflows based on goals rather than fixed scripts.

4 maturity levels:

LevelApproachWhat AI DoesHuman Role
L1AI-AssistedGenerate test cases, suggest assertions, auto-complete scriptsReviews everything
L2AI-AugmentedAuto-generate regression suites, self-heal broken selectors, generate test dataFocuses on exploratory testing
L3Agentic QADecides what to test based on code changes, risk, and coverage gapsSets goals, reviews results
L4Self-HealingDetects failures, diagnoses root cause, fixes tests, re-runs, reportsHandles only novel failures

Agentic QA workflow patterns:

PatternQA ApplicationValue
ChainingRequirements โ†’ Generate test cases โ†’ Prioritize by risk โ†’ Generate test data โ†’ Output suite6 hours manual โ†’ 30 min review
RoutingCode change โ†’ Classify (UI/API/DB) โ†’ Route to targeted test suite โ†’ Run โ†’ Report4-hour full regression โ†’ 20-min targeted
OrchestrationAgent explores app โ†’ Detects anomalies โ†’ Generates bug reports โ†’ Prioritizes โ†’ EscalatesFinds edge cases humans miss; runs 24/7
ParallelizationFailed tests โ†’ Check UI change? + Check API change? + Real bug? โ†’ Auto-fix or file bug40% reduction in test maintenance
For AnyCompany QA team: Start at L2 (AI-Augmented) โ€” use AI to generate test cases from requirements and auto-generate regression suites. High-value starting points: regulatory test coverage from MAS circulars, claims processing regression, policy document validation after system updates.

Agentic AI Foundation โ€” Key Takeaways

  • Agents โ‰  Chatbots: Agents observe, plan, act, and reflect in a loop. They use tools, access data, and make decisions within guardrails.
  • 4 Workflow Patterns: Chaining (sequential), Parallelization (simultaneous), Routing (classify & direct), Orchestration (dynamic + human-in-the-loop)
  • Design before build: The Agent Design Canvas captures what to build, how it should behave, and when humans take over โ€” before any code is written.
  • Guardrails are non-negotiable: Every agent needs clear boundaries โ€” what it must NOT do, when it must escalate, and what data it can access.
  • Start small, prove value: Pick one workflow, build one agent, measure the impact. Then expand.
๐ŸŽฏ Your deliverable: The Agent Design Canvas you created today is the document you hand to your tech team. It tells them exactly what to build, what guardrails to enforce, and how to measure success. Start the conversation this week.