INTERACTIVE EXPLAINER
The techniques that turn vague AI outputs into production-grade, auditable results โ with interactive before/after comparisons and a live prompt builder.
80% of prompt quality comes from 4 fundamentals. Master these and every prompt you write โ whether it's a claims summary, underwriting memo, or compliance check โ will be dramatically better.
Say exactly what you mean. If a colleague would ask "what do you mean?" โ your prompt needs work.
Give the AI the background it needs: domain, data, situation, constraints. Without context, it guesses.
Tell the AI who to be. A "claims adjudicator" focuses on different signals than a "customer service agent."
Define what "done" looks like โ format, length, structure, style. No framing = unpredictable output.
Context is the most impactful pillar for insurance. Skip any type and the output suffers in a specific way:
| Type | What it tells the AI | If you skip it... | Insurance example |
|---|---|---|---|
| Domain | Industry, market, business area | Generic, non-specific answers | "In the context of Singapore life insurance underwriting..." |
| Data | Specific numbers, records, documents | AI hallucinates plausible numbers | "Here is the policyholder's claim history and medical report: [data]" |
| Situational | Why you need this now โ the trigger | Wrong tone, urgency, framing | "Preparing for MAS regulatory submission" / "Flagged by fraud detection" |
| Constraints | Rules, limits, requirements | Ignores your standards | "All amounts in SGD" / "Follow MAS Notice 127 guidelines" |
Prompt: "Review this claim"
The claim seems to be legitimate. The policyholder has been with us for a while and the medical documentation looks reasonable. There are some inconsistencies but that's fairly normal. Overall I would recommend processing the claim.
โ ๏ธ Vague, no data cited, no structure, no decision rating, not actionable
Prompt: Role + Context + Structure + Constraints
DECISION: ๐ก REFER TO SENIOR ADJUDICATOR
Claim Analysis: Critical illness claim for $200K. Diagnosis date 14 months post-inception โ within 24-month contestability window.
Red Flag: Pre-existing condition not disclosed on application. GP records show consultation 3 months before policy start.
Action: 1) Request full medical history from attending physician (Claims Ops, 5 days) 2) Refer to Special Investigations Unit (SIU, immediate)
โ Structured, data-cited, specific actions with owners and timelines
Insurance decisions require multi-step logic. CoT makes reasoning visible and auditable โ the AI shows its work, which is critical for regulatory compliance and dispute resolution.
Just add "Think step by step." No examples needed. Best for quick calculations and simple logic.
Provide one example with reasoning. The AI follows the exact same pattern. Best for consistent processes.
"First identify key factors, then analyze." Forces prioritization before writing. Best for complex analysis.
"Solve 3 ways, report majority." Multiple approaches catch what a single analysis misses. Best for high-stakes.
Question: "Should we approve this $500K term life application?"
Yes, the application looks acceptable. The applicant has a reasonable income and no major health issues reported. I would recommend approval at standard rates.
โ ๏ธ No math. No reasoning. Not auditable. Could miss critical risk factors.
Annual income: $180K SGD
โ Sum assured requested: $500K
= Income multiple: 2.8x (within 10x guideline โ
)
Age: 42, non-smoker, BMI 24.2
= Mortality class: Preferred
Existing coverage: $200K (other insurer)
= Total exposure: $700K = 3.9x income
โ
Within financial underwriting limits โ APPROVE at Preferred rates
โ Every step visible. Math auditable. Conclusion backed by evidence.
For high-stakes decisions, analyze from 3 independent angles and take the majority vote:
| Approach | Analysis | Conclusion |
|---|---|---|
| 1. Medical | Hospitalization claim 3 days after policy inception. Diagnosis: acute appendicitis (sudden onset, not pre-existing). | ๐ข LEGITIMATE |
| 2. Financial | Claim amount $45K for 3-day stay at private hospital. Itemized bill matches typical appendectomy costs in Singapore. | ๐ข LEGITIMATE |
| 3. Behavioral | Policyholder upgraded plan from basic to comprehensive 2 weeks before admission. No prior claims in 5 years on basic plan. | ๐ด SUSPICIOUS |
Majority: 2/3 LEGITIMATE. Approach 3 alone would have flagged this for investigation. The majority vote provides balanced judgment โ but the timing anomaly should still be noted in the file.
Same data, dramatically different insights โ just by changing who the AI "is." The AI was trained on millions of documents written by different professionals. When you assign a persona, you activate that specific knowledge cluster.
Data: Critical illness claim for $300K. Policyholder diagnosed with Stage 2 cancer, 18 months after policy inception. Non-disclosure of family history on application.
RECOMMENDATION: DECLINE โ Material Non-Disclosure
Family history of cancer (father, uncle) not disclosed on application. Within 24-month contestability period. Had applicant disclosed, underwriting would have required additional screening and likely applied exclusion or loading.
Action: Issue decline letter citing Section 25 Insurance Act. Offer ex-gratia refund of premiums paid. Prepare for potential FIDReC dispute.
RECOMMENDATION: PARTIAL APPROVAL โ Fair Dealing Approach
Policyholder is a loyal customer (3 other policies, 8 years). Non-disclosure appears inadvertent โ family history question was buried in page 12 of a 15-page form. MAS Fair Dealing Guidelines emphasize proportionate response.
Action: Approve at reduced sum ($200K, reflecting what would have been offered with disclosure). Retain customer relationship. Document reasoning for audit trail.
Both are valid. The investigator sees contractual breach. The CX director sees customer retention and fair dealing. Neither is wrong โ they serve different business objectives.
Get 3 perspectives in one prompt โ no need to schedule 3 meetings:
| Perspective | Focus | Key finding |
|---|---|---|
| ๐ก๏ธ Chief Actuary | Loss ratios, reserves, pricing adequacy | "Approving borderline claims increases loss ratio by 2.3 points" |
| ๐ Head of Distribution | Persistency, NPS, agent retention | "Claim disputes are #1 driver of policy lapses and agent attrition" |
| โ๏ธ Chief Compliance Officer | MAS guidelines, PDPA, fair dealing | "MAS Fair Dealing Outcome 5 requires claims to be paid fairly and promptly" |
Synthesis: Approve claim with enhanced documentation. Update underwriting questionnaire to make family history more prominent. Monitor loss ratio impact quarterly. Report to Board Risk Committee if loss ratio exceeds 65%.
Multi-agent framing is a well-established prompt engineering technique with several names in the research literature:
| Technique | Source | Key idea |
|---|---|---|
| Solo Performance Prompting (SPP) | Wang et al., 2023 | A single LLM simulates multiple personas that collaborate internally โ "cognitive synergy through multi-persona self-collaboration" |
| Multi-Persona Thinking (MPT) | arXiv 2025 | Dialectical reasoning from multiple perspectives to reduce bias and improve decision quality |
| Town Hall Debate Prompting | arXiv 2025 | Splices a language model into multiple personas that debate one another to reach a conclusion |
| Self-Consistency | Wang et al., 2022 | Generate multiple reasoning paths and aggregate โ the broader technique family that multi-perspective builds on |
Consistent format + grounded in YOUR data = production-safe outputs.
Different every time. Hard to compare. Can't feed into systems. Requires human parsing.
Consistent format. Comparable across claims. Machine-parseable. Scannable by busy stakeholders.
Tell the AI exactly what shape the output should take. The more specific your format instructions, the more consistent the results.
| Technique | Prompt example | What you get |
|---|---|---|
| Named sections | "Use these sections: Summary, Risk Factors, Recommendation" | Same headings every time โ scannable, comparable |
| Table format | "Present as a table: Metric | Value | Benchmark | Status" | Aligned data, easy to paste into Excel |
| JSON output | "Return JSON: {decision, confidence, reasoning, actions[]}" | Machine-readable, feeds into dashboards or APIs |
| Numbered actions | "List 3 actions. Each: action, owner, deadline, priority (H/M/L)" | Actionable items with accountability |
| Rating + justification | "Give an APPROVE/REFER/DECLINE decision. Justify in exactly 2 sentences." | Consistent decision format across all reviews |
| Length control | "Executive summary: max 3 sentences. Detail: max 200 words." | Right depth for the audience |
When you ask AI to produce a report, analysis, or any reusable document โ ask for Markdown. It's the format that works best for both humans and AI.
| Format | Human readable | AI readable | Token cost | Reusable |
|---|---|---|---|---|
| โ | โ Can't parse | N/A | โ | |
| Word (.docx) | โ | โ ๏ธ Partial | N/A | โ |
| HTML | โ ๏ธ Tags clutter | โ | High (~20 tokens/heading) | โ |
| Markdown โ | โ | โ | Low (~8 tokens/heading) | โ |
How to ask for it:
.kiro/steering/rules.md), skills (SKILL.md), agent configs โ is Markdown. It's the interface layer between you and AI: structured enough for machines, readable enough for humans, and 60% fewer tokens than HTML.This isn't just a convention โ research and industry practice back it up:
| Finding | Impact | Source |
|---|---|---|
| Markdown vs HTML token usage | 60% fewer tokens for same content structure | Token comparison (heading: ~8 vs ~20 tokens) |
| Markdown vs JSON for LLM comprehension | 16% average token savings with equal or better accuracy | Format performance benchmarks |
| Table extraction accuracy | Markdown 60.7% vs HTML 53.6% | ReleasePad, 2025 |
| RAG retrieval with clean Markdown | Up to 35% better retrieval accuracy, 20-30% fewer tokens | AnythingMD |
| llms.txt web standard (Sept 2024) | Websites now serve Markdown specifically for AI agents | Jeremy Howard, Answer.AI |
| LLM Markdown awareness research | LLMs are expected to produce structured Markdown for readability | arXiv:2501.15000, 2025 |
llms.txt standard (proposed by Jeremy Howard of fast.ai in September 2024) is like robots.txt but for AI โ websites now serve Markdown files at their root specifically for AI agents to read. When you write a steering file, a SKILL.md, or ask for a report โ Markdown is the right default.This section is for technical team members (Cloud COE, Platforms). Most business users can skip this โ the plain-text techniques above are all you need for daily use.
When building prompt templates at the code level (Bedrock API, application backends), developers often wrap prompt sections in XML tags. This is how Anthropic recommends structuring complex API calls โ the tags create unambiguous boundaries between instructions, data, and constraints.
| Pattern | Tags | When to use |
|---|---|---|
| Claims analysis | <data> <task> <format> | Analyzing claims, medical reports, loss assessments |
| Policy Q&A | <document> <question> <rules> | Policy lookups, coverage checks, exclusion queries |
| Multi-step | <context> <step1> <step2> | Complex workflows, underwriting assessments |
| Review | <draft> <criteria> <instructions> | Reviewing reports, correspondence, regulatory submissions |
<data>...</data> tags behind the scenes before sending to the Bedrock API. You design the content; your tech team handles the XML structure. For daily use in Claude or Kiro, plain-text headers (### ROLE, ### CONTEXT) work just as well.Without grounding rules, the AI mixes its training data with your documents โ you can't trace which claims come from where. In an audit or MAS inspection, that's a disaster.
Q: "What happens if a policyholder misses 3 premium payments?"
The policy typically enters a grace period and then lapses. The policyholder usually has 30 days to pay before the policy is terminated. They may be able to reinstate within 2 years with evidence of insurability.
โ ๏ธ Sounds plausible but NONE of this may match your actual policy wording. "Typically" and "usually" = hallucination.
Q: Same question + policy document attached + grounding rules
Per AnyCompany Life Policy Terms v4.1:
โข After 1 missed payment โ 31-day grace period [Section 12.1]
โข After grace period โ Automatic Premium Loan if cash value sufficient [Section 12.3]
โข If no cash value โ Policy lapses [Section 12.4]
โข Reinstatement: within 24 months, subject to health declaration [Section 13.1]
โข [INSUFFICIENT DATA: No information on premium waiver rider in provided document]
โ Every claim cites a section. Admits what it doesn't know. No hallucination.
Toggle techniques on/off to see how the prompt AND the AI's response change. Watch quality improve as you add each technique.
Recognize these patterns? Fix them with one-line additions to your prompt.
| Mistake | Why it hurts | Quick fix |
|---|---|---|
| ๐ณ The Kitchen Sink | Cramming 5 tasks into 1 prompt | One task per prompt, chain results |
| ๐ The Blank Canvas | No examples = AI guesses your format | Show 1-2 examples of desired output |
| ๐ The Trust Fall | No grounding = confident hallucinations | "ONLY from provided data" |
| ๐ The Vague Ask | "Analyze this" โ analyze what, how, for whom? | Specify audience, format, length |
| โฑ๏ธ The One-Shot Wonder | Expecting perfection on first try | Plan for 2-3 refinement turns |
| ๐ The Copy-Paste Trap | Same prompt for different models | Tune syntax per model family |
| โ๏ธ The Set-and-Forget | Never re-testing after model updates | Monthly prompt health checks |
Every production-quality prompt goes through this cycle:
| Round | What you do | Result |
|---|---|---|
| 1. Baseline | Write prompt using 4 pillars. Run 3 times. | See what AI gets right and wrong (~60% quality) |
| 2. Fix failures | Add negative constraints + example of good output. Run 3 more. | Consistency jumps to ~85% |
| 3. Polish | Add self-review step. Tighten format. Test edge cases. | Production-ready at ~95% |
Negative constraints prevent common failure modes:
| Problem | Add this constraint |
|---|---|
| AI adds unsolicited opinions | "Do not include personal opinions or speculation" |
| AI uses data not in your input | "Do not reference any data outside the provided documents" |
| AI writes too much | "Do not exceed 300 words" |
| AI hedges everything | "Do not use phrases like 'it depends' or 'generally speaking'" |
| AI explains obvious things | "Do not explain what life insurance is or how premiums work" |
| AI invents numbers | "If a metric is not in the data, write [DATA NOT AVAILABLE]" |
It doesn't. Each session is completely isolated. The AI has zero memory of previous conversations.
| What persists | What doesn't |
|---|---|
| โ Files in your workspace (reports, templates, code) | โ Chat conversation history |
| โ Steering files (.kiro/steering/) โ loaded every session | โ What you said 3 sessions ago |
| โ Skills (.kiro/skills/) โ activated by keywords | โ Old tabs or closed sessions |
| โ Custom agents (.kiro/agents/) โ invoked by name | โ Your "relationship" with the AI |