INTERACTIVE EXPLAINER

Prompt Engineering Best Practices

The techniques that turn vague AI outputs into production-grade, auditable results โ€” with interactive before/after comparisons and a live prompt builder.

๐Ÿ“– Reference โšก Interactive ๐Ÿฅ Insurance Examples

๐ŸŽฏ The 4 Pillars of Effective Prompts

80% of prompt quality comes from 4 fundamentals. Master these and every prompt you write โ€” whether it's a claims summary, underwriting memo, or compliance check โ€” will be dramatically better.

๐ŸŽฏ

1. Clarity

Say exactly what you mean. If a colleague would ask "what do you mean?" โ€” your prompt needs work.

๐Ÿ“š

2. Context

Give the AI the background it needs: domain, data, situation, constraints. Without context, it guesses.

๐Ÿ‘ค

3. Role Assignment

Tell the AI who to be. A "claims adjudicator" focuses on different signals than a "customer service agent."

๐Ÿ“‹

4. Output Framing

Define what "done" looks like โ€” format, length, structure, style. No framing = unpredictable output.

๐Ÿ“š 4 Types of Context

Context is the most impactful pillar for insurance. Skip any type and the output suffers in a specific way:

TypeWhat it tells the AIIf you skip it...Insurance example
DomainIndustry, market, business areaGeneric, non-specific answers"In the context of Singapore life insurance underwriting..."
DataSpecific numbers, records, documentsAI hallucinates plausible numbers"Here is the policyholder's claim history and medical report: [data]"
SituationalWhy you need this now โ€” the triggerWrong tone, urgency, framing"Preparing for MAS regulatory submission" / "Flagged by fraud detection"
ConstraintsRules, limits, requirementsIgnores your standards"All amounts in SGD" / "Follow MAS Notice 127 guidelines"

๐Ÿ‘€ See the Difference

โŒ Without 4 Pillars

Prompt: "Review this claim"

The claim seems to be legitimate. The policyholder has been with us for a while and the medical documentation looks reasonable. There are some inconsistencies but that's fairly normal. Overall I would recommend processing the claim.

โš ๏ธ Vague, no data cited, no structure, no decision rating, not actionable

โœ… With 4 Pillars

Prompt: Role + Context + Structure + Constraints

DECISION: ๐ŸŸก REFER TO SENIOR ADJUDICATOR
Claim Analysis: Critical illness claim for $200K. Diagnosis date 14 months post-inception โ€” within 24-month contestability window.
Red Flag: Pre-existing condition not disclosed on application. GP records show consultation 3 months before policy start.
Action: 1) Request full medical history from attending physician (Claims Ops, 5 days) 2) Refer to Special Investigations Unit (SIU, immediate)

โœ“ Structured, data-cited, specific actions with owners and timelines

๐Ÿ’ก
Same data, dramatically different quality. The difference is entirely in the prompt โ€” not the AI model.

๐Ÿง  Chain-of-Thought Reasoning

Insurance decisions require multi-step logic. CoT makes reasoning visible and auditable โ€” the AI shows its work, which is critical for regulatory compliance and dispute resolution.

๐Ÿ’ฌ

Zero-Shot CoT

Just add "Think step by step." No examples needed. Best for quick calculations and simple logic.

๐Ÿ“

Few-Shot CoT

Provide one example with reasoning. The AI follows the exact same pattern. Best for consistent processes.

๐Ÿ”ญ

Step-Back

"First identify key factors, then analyze." Forces prioritization before writing. Best for complex analysis.

๐Ÿ—ณ๏ธ

Self-Consistency

"Solve 3 ways, report majority." Multiple approaches catch what a single analysis misses. Best for high-stakes.

๐Ÿ‘€ Zero-Shot CoT: Before & After

Question: "Should we approve this $500K term life application?"

โŒ Without CoT

Yes, the application looks acceptable. The applicant has a reasonable income and no major health issues reported. I would recommend approval at standard rates.

โš ๏ธ No math. No reasoning. Not auditable. Could miss critical risk factors.

โœ… With "Think step by step"

Annual income: $180K SGD
โˆ’ Sum assured requested: $500K
= Income multiple: 2.8x (within 10x guideline โœ…)
Age: 42, non-smoker, BMI 24.2
= Mortality class: Preferred
Existing coverage: $200K (other insurer)
= Total exposure: $700K = 3.9x income
โœ… Within financial underwriting limits โ€” APPROVE at Preferred rates

โœ“ Every step visible. Math auditable. Conclusion backed by evidence.

๐Ÿ—ณ๏ธ Self-Consistency: 3 Approaches, 1 Answer

For high-stakes decisions, analyze from 3 independent angles and take the majority vote:

ApproachAnalysisConclusion
1. MedicalHospitalization claim 3 days after policy inception. Diagnosis: acute appendicitis (sudden onset, not pre-existing).๐ŸŸข LEGITIMATE
2. FinancialClaim amount $45K for 3-day stay at private hospital. Itemized bill matches typical appendectomy costs in Singapore.๐ŸŸข LEGITIMATE
3. BehavioralPolicyholder upgraded plan from basic to comprehensive 2 weeks before admission. No prior claims in 5 years on basic plan.๐Ÿ”ด SUSPICIOUS

Majority: 2/3 LEGITIMATE. Approach 3 alone would have flagged this for investigation. The majority vote provides balanced judgment โ€” but the timing anomaly should still be noted in the file.

โš ๏ธ
Insurance rule: Any decision that could be audited by MAS or challenged in a dispute should use CoT โ€” the reasoning trail IS your documentation.

๐Ÿ‘ค Role & Persona Prompting

Same data, dramatically different insights โ€” just by changing who the AI "is." The AI was trained on millions of documents written by different professionals. When you assign a persona, you activate that specific knowledge cluster.

The Persona Formula

You are [TITLE] at [COMPANY TYPE]
with [X years] of experience in [SPECIALTY].
You are known for [CHARACTERISTIC].
When [SITUATION], you always [BEHAVIOR].
๐Ÿ’ก
The last two fields โ€” CHARACTERISTIC and BEHAVIOR โ€” matter most. "Conservative" vs "growth-focused" produces completely different recommendations from the same data.

๐Ÿ‘€ Same Claim, Different Eyes

Data: Critical illness claim for $300K. Policyholder diagnosed with Stage 2 cancer, 18 months after policy inception. Non-disclosure of family history on application.

๐Ÿ›ก๏ธ Claims Investigation Lead

RECOMMENDATION: DECLINE โ€” Material Non-Disclosure

Family history of cancer (father, uncle) not disclosed on application. Within 24-month contestability period. Had applicant disclosed, underwriting would have required additional screening and likely applied exclusion or loading.

Action: Issue decline letter citing Section 25 Insurance Act. Offer ex-gratia refund of premiums paid. Prepare for potential FIDReC dispute.

๐Ÿ’š Customer Experience Director

RECOMMENDATION: PARTIAL APPROVAL โ€” Fair Dealing Approach

Policyholder is a loyal customer (3 other policies, 8 years). Non-disclosure appears inadvertent โ€” family history question was buried in page 12 of a 15-page form. MAS Fair Dealing Guidelines emphasize proportionate response.

Action: Approve at reduced sum ($200K, reflecting what would have been offered with disclosure). Retain customer relationship. Document reasoning for audit trail.

Both are valid. The investigator sees contractual breach. The CX director sees customer retention and fair dealing. Neither is wrong โ€” they serve different business objectives.

๐Ÿค Multi-Agent Framing

Get 3 perspectives in one prompt โ€” no need to schedule 3 meetings:

PerspectiveFocusKey finding
๐Ÿ›ก๏ธ Chief ActuaryLoss ratios, reserves, pricing adequacy"Approving borderline claims increases loss ratio by 2.3 points"
๐Ÿ“Š Head of DistributionPersistency, NPS, agent retention"Claim disputes are #1 driver of policy lapses and agent attrition"
โš–๏ธ Chief Compliance OfficerMAS guidelines, PDPA, fair dealing"MAS Fair Dealing Outcome 5 requires claims to be paid fairly and promptly"

Synthesis: Approve claim with enhanced documentation. Update underwriting questionnaire to make family history more prominent. Monitor loss ratio impact quarterly. Report to Board Risk Committee if loss ratio exceeds 65%.

๐Ÿ”
The synthesis is where the real insight lives. No single perspective dominates โ€” the balanced recommendation is stronger than any individual view.

The Research Behind It

Multi-agent framing is a well-established prompt engineering technique with several names in the research literature:

TechniqueSourceKey idea
Solo Performance Prompting (SPP)Wang et al., 2023A single LLM simulates multiple personas that collaborate internally โ€” "cognitive synergy through multi-persona self-collaboration"
Multi-Persona Thinking (MPT)arXiv 2025Dialectical reasoning from multiple perspectives to reduce bias and improve decision quality
Town Hall Debate PromptingarXiv 2025Splices a language model into multiple personas that debate one another to reach a conclusion
Self-ConsistencyWang et al., 2022Generate multiple reasoning paths and aggregate โ€” the broader technique family that multi-perspective builds on
๐Ÿ’ก
Why it works: LLMs are trained on millions of documents written by different professionals. When you assign a persona, you activate that specific knowledge cluster. Asking for 3 personas in one prompt triggers 3 distinct "knowledge activations" โ€” producing genuinely different analyses, not just rephrased versions of the same answer.
๐Ÿ”—
Agentic AI connection: Here, you simulate multiple perspectives in a single prompt. In the agentic world, you'll see the automated version โ€” the Parallelization pattern โ€” where each perspective actually runs as a separate AI agent simultaneously, and a real aggregator combines the results. Same concept, automated at scale.

๐Ÿ“‹ Structured Outputs & RAG Grounding

Consistent format + grounded in YOUR data = production-safe outputs.

Why Structure Matters

โŒ Unstructured = Conversation

Different every time. Hard to compare. Can't feed into systems. Requires human parsing.

โœ… Structured = Form

Consistent format. Comparable across claims. Machine-parseable. Scannable by busy stakeholders.

How to Prompt for Structured Output

Tell the AI exactly what shape the output should take. The more specific your format instructions, the more consistent the results.

TechniquePrompt exampleWhat you get
Named sections"Use these sections: Summary, Risk Factors, Recommendation"Same headings every time โ€” scannable, comparable
Table format"Present as a table: Metric | Value | Benchmark | Status"Aligned data, easy to paste into Excel
JSON output"Return JSON: {decision, confidence, reasoning, actions[]}"Machine-readable, feeds into dashboards or APIs
Numbered actions"List 3 actions. Each: action, owner, deadline, priority (H/M/L)"Actionable items with accountability
Rating + justification"Give an APPROVE/REFER/DECLINE decision. Justify in exactly 2 sentences."Consistent decision format across all reviews
Length control"Executive summary: max 3 sentences. Detail: max 200 words."Right depth for the audience

Full Example: Combining Techniques

OUTPUT FORMAT: 1. Decision โ€” APPROVE/REFER/DECLINE with 2-sentence justification 2. Key Metrics Table: | Factor | Finding | Benchmark | Assessment | 3. Analysis โ€” max 150 words, cite specific policy sections 4. Recommended Actions: - Numbered list, each with: action, owner, deadline 5. JSON Summary (for system integration): {"decision": "...", "confidence": 0-100, "top_risk": "..."}
๐Ÿ’ก
Pro tip: You can mix human-readable sections (1-4) with machine-readable JSON (5) in the same prompt. The AI handles both formats in one response. This is how production templates work โ€” the human reads the narrative, the system reads the JSON.

The Best Default Format: Markdown (.md)

When you ask AI to produce a report, analysis, or any reusable document โ€” ask for Markdown. It's the format that works best for both humans and AI.

FormatHuman readableAI readableToken costReusable
PDFโœ…โŒ Can't parseN/AโŒ
Word (.docx)โœ…โš ๏ธ PartialN/AโŒ
HTMLโš ๏ธ Tags clutterโœ…High (~20 tokens/heading)โœ…
Markdown โœ“โœ…โœ…Low (~8 tokens/heading)โœ…

How to ask for it:

Save the output as "claims-review.md" with: - ## headings for each section - | tables | for data comparisons - - bullet lists for action items
๐Ÿ”—
Why this matters for you: In the agentic AI world, every artifact you create โ€” steering files (.kiro/steering/rules.md), skills (SKILL.md), agent configs โ€” is Markdown. It's the interface layer between you and AI: structured enough for machines, readable enough for humans, and 60% fewer tokens than HTML.

The Research Behind Markdown for AI

This isn't just a convention โ€” research and industry practice back it up:

FindingImpactSource
Markdown vs HTML token usage60% fewer tokens for same content structureToken comparison (heading: ~8 vs ~20 tokens)
Markdown vs JSON for LLM comprehension16% average token savings with equal or better accuracyFormat performance benchmarks
Table extraction accuracyMarkdown 60.7% vs HTML 53.6%ReleasePad, 2025
RAG retrieval with clean MarkdownUp to 35% better retrieval accuracy, 20-30% fewer tokensAnythingMD
llms.txt web standard (Sept 2024)Websites now serve Markdown specifically for AI agentsJeremy Howard, Answer.AI
LLM Markdown awareness researchLLMs are expected to produce structured Markdown for readabilityarXiv:2501.15000, 2025
๐Ÿ’ก
The industry is converging on Markdown as the standard interface between humans and AI. LLMs are trained on it, tools expect it, and it costs less. The llms.txt standard (proposed by Jeremy Howard of fast.ai in September 2024) is like robots.txt but for AI โ€” websites now serve Markdown files at their root specifically for AI agents to read. When you write a steering file, a SKILL.md, or ask for a report โ€” Markdown is the right default.

๐Ÿ”ง Advanced: XML Tags for Claude (Optional)

This section is for technical team members (Cloud COE, Platforms). Most business users can skip this โ€” the plain-text techniques above are all you need for daily use.

When building prompt templates at the code level (Bedrock API, application backends), developers often wrap prompt sections in XML tags. This is how Anthropic recommends structuring complex API calls โ€” the tags create unambiguous boundaries between instructions, data, and constraints.

// Typically constructed in application code, not typed by hand: <role>Senior Claims Adjudicator, 12 years in life insurance</role> <data> Policyholder: John Tan (Policy #LI-2024-88421) Claim type: Critical Illness Sum assured: $300,000 SGD Diagnosis: Stage 2 colorectal cancer Policy inception: 18 months ago </data> <task>Assess claim validity. Produce an APPROVE/REFER/DECLINE decision.</task> <constraints> - ONLY use data in <data> tags - Max 300 words - Cite relevant policy sections </constraints>
PatternTagsWhen to use
Claims analysis<data> <task> <format>Analyzing claims, medical reports, loss assessments
Policy Q&A<document> <question> <rules>Policy lookups, coverage checks, exclusion queries
Multi-step<context> <step1> <step2>Complex workflows, underwriting assessments
Review<draft> <criteria> <instructions>Reviewing reports, correspondence, regulatory submissions
๐Ÿ”
Where XML tags live in practice: In production systems, the developer builds the prompt template with XML tags programmatically. The end user fills in a form or pastes data โ€” the application wraps it in <data>...</data> tags behind the scenes before sending to the Bedrock API. You design the content; your tech team handles the XML structure. For daily use in Claude or Kiro, plain-text headers (### ROLE, ### CONTEXT) work just as well.

๐Ÿ”’ RAG Grounding: Only YOUR Data

Without grounding rules, the AI mixes its training data with your documents โ€” you can't trace which claims come from where. In an audit or MAS inspection, that's a disaster.

CRITICAL GROUNDING RULES: 1. Base your answer ONLY on the provided documents 2. After each claim, cite: [Doc Name, Section] 3. If not in documents: "Not available in provided documents" 4. Do NOT use outside knowledge
โŒ Without RAG Grounding

Q: "What happens if a policyholder misses 3 premium payments?"

The policy typically enters a grace period and then lapses. The policyholder usually has 30 days to pay before the policy is terminated. They may be able to reinstate within 2 years with evidence of insurability.

โš ๏ธ Sounds plausible but NONE of this may match your actual policy wording. "Typically" and "usually" = hallucination.

โœ… With RAG Grounding

Q: Same question + policy document attached + grounding rules

Per AnyCompany Life Policy Terms v4.1:
โ€ข After 1 missed payment โ†’ 31-day grace period [Section 12.1]
โ€ข After grace period โ†’ Automatic Premium Loan if cash value sufficient [Section 12.3]
โ€ข If no cash value โ†’ Policy lapses [Section 12.4]
โ€ข Reinstatement: within 24 months, subject to health declaration [Section 13.1]
โ€ข [INSUFFICIENT DATA: No information on premium waiver rider in provided document]

โœ“ Every claim cites a section. Admits what it doesn't know. No hallucination.

๐Ÿ”ง Interactive Prompt Builder

Toggle techniques on/off to see how the prompt AND the AI's response change. Watch quality improve as you add each technique.

๐Ÿ“ Your Prompt 0 words
Loading...
๐Ÿค– AI Response
Loading...
๐Ÿ’ก What changed: Toggle techniques above to see how the AI response improves.

๐Ÿ“Š Quality Score

Completeness
2/5
Data Grounding
1/5
Actionability
1/5
Consistency
2/5
6/20
Needs work โ€” toggle more techniques

๐Ÿ” Issues in AI Response

โš ๏ธ 7 Prompt Mistakes Everyone Makes

Recognize these patterns? Fix them with one-line additions to your prompt.

MistakeWhy it hurtsQuick fix
๐Ÿณ The Kitchen SinkCramming 5 tasks into 1 promptOne task per prompt, chain results
๐Ÿ“„ The Blank CanvasNo examples = AI guesses your formatShow 1-2 examples of desired output
๐Ÿ™ˆ The Trust FallNo grounding = confident hallucinations"ONLY from provided data"
๐Ÿ” The Vague Ask"Analyze this" โ€” analyze what, how, for whom?Specify audience, format, length
โฑ๏ธ The One-Shot WonderExpecting perfection on first tryPlan for 2-3 refinement turns
๐Ÿ“‹ The Copy-Paste TrapSame prompt for different modelsTune syntax per model family
โš™๏ธ The Set-and-ForgetNever re-testing after model updatesMonthly prompt health checks

๐Ÿ”„ The 3-Round Improvement Workflow

Every production-quality prompt goes through this cycle:

RoundWhat you doResult
1. BaselineWrite prompt using 4 pillars. Run 3 times.See what AI gets right and wrong (~60% quality)
2. Fix failuresAdd negative constraints + example of good output. Run 3 more.Consistency jumps to ~85%
3. PolishAdd self-review step. Tighten format. Test edge cases.Production-ready at ~95%
๐Ÿ’ก
Total time: 15-20 minutes to go from first draft to production template. That template then saves hours every week โ€” imagine every claims summary, underwriting memo, or compliance check following the same quality standard automatically.

๐Ÿšซ Tell the AI What NOT to Do

Negative constraints prevent common failure modes:

ProblemAdd this constraint
AI adds unsolicited opinions"Do not include personal opinions or speculation"
AI uses data not in your input"Do not reference any data outside the provided documents"
AI writes too much"Do not exceed 300 words"
AI hedges everything"Do not use phrases like 'it depends' or 'generally speaking'"
AI explains obvious things"Do not explain what life insurance is or how premiums work"
AI invents numbers"If a metric is not in the data, write [DATA NOT AVAILABLE]"
๐Ÿ”
Source: Claude's prompting best practices recommend telling Claude what to do instead of what not to do for general instructions, but negative constraints are highly effective for preventing specific failure modes โ€” especially in insurance where hallucinated policy terms or fabricated claim amounts are dangerous. Claude Prompting Best Practices โ†’

๐Ÿง  The #1 Misconception: "AI Remembers Me"

It doesn't. Each session is completely isolated. The AI has zero memory of previous conversations.

โŒ What people think
  • "It remembers our conversation from last week"
  • "I should keep this tab open so it doesn't forget"
  • "My old sessions are giving it context"
โœ… How it actually works
  • Each session starts with zero memory
  • Old tabs have no effect on new sessions
  • Closing old sessions is safe โ€” cosmetic, not functional
What persistsWhat doesn't
โœ… Files in your workspace (reports, templates, code)โŒ Chat conversation history
โœ… Steering files (.kiro/steering/) โ€” loaded every sessionโŒ What you said 3 sessions ago
โœ… Skills (.kiro/skills/) โ€” activated by keywordsโŒ Old tabs or closed sessions
โœ… Custom agents (.kiro/agents/) โ€” invoked by nameโŒ Your "relationship" with the AI
๐Ÿ’ก
The mental model: chat is ephemeral, files are permanent. Save important outputs as files. Reference files (not old chats) when you need context in a new session. Steering files and skills ARE the AI's persistent memory โ€” they're loaded automatically into every new session.