Prompt Best Practices
A concrete translation into prompt design.
What people say is already enough. A good prompt helps AI make it visible, without adding anything to it.
- Dembrane
- Maarten Essenburg
- the Doesburg engagement
- the Field Guide
A reference for anyone designing prompts for their own sessions that follow this philosophy. Not generic "prompt engineering 101", but philosophy translated into concrete design rules. The seven baseline guidelines for public AI output below are always active; later on we refer to "guideline #N" to recall the same rule quickly.
Living document, continuously evolving. First version autumn 2025. Published publicly 2026-05-20. Last update 2026-05-19.
Part of Thoughtful Social AI. License: CC-BY-SA 4.0. View raw markdown for copy-paste into your AI.
What this document references. As you read, you'll come across a number of specific tools, projects, and people that are central to the author's practice. The principles and techniques are generally applicable; these concrete references serve as context, not as a requirement.
- Dembrane: a transcript and dialogue platform for group work where many of these prompts were developed
- Maarten Essenburg: facilitator colleague (one of the practice examples in this doc)
- the Doesburg engagement: a bottom-up community engagement around a caring community (the reference source for the frustration-as-fuel and ownership examples)
- the Field Guide: Social AI Field Guide at socialaiveldgids.nl (the practice context these prompts land in)
What people say is already enough. A good prompt helps AI make that visible, without adding anything to it.
A reference for anyone designing prompts for their own sessions that follow this philosophy. Not a generic "prompt engineering 101", a philosophy translated into concrete design rules. The seven baseline guidelines for public AI output below are always-active; further on we refer to "guideline #N" to quickly recall the same rule.
Seven baseline guidelines for public AI output
These seven guidelines form the always-active foundation under every prompt whose output leaves the workspace (participant output, public posts, client reports, AI output that reaches participants through a platform). Further on in this document we refer to them in short as "guideline #N".
| # | Guideline | One-liner |
|---|---|---|
| 1 | Em-dash ban for outward output | No โ or --. Comma, colon, parentheses, or a new sentence. Internal (chat, working documents) is fine. |
| 2 | AI-protagonist ban (with inference exception) | No "AI noticed" / "our analysis" / "the algorithm thinks". Exception: explicit inference as (a) an open question to the group, (b) marked without an ownership claim, (c) valuable. DO: "In the analysis it stood out that... is this a shared concern?". DON'T: "Our analysis shows that...". |
| 3 | Privacy and names: simplest human form | Default = simplest anonymized: "a participant" / "participant 1". Specificity only when it adds value. Hierarchy: (1) "a participant", (2) role description, (3) workstream or group tag, (4) full name only with permission. Redaction placeholders never visible. |
| 4 | Recognition test | Test after output: "yes, that's what we said" or "sounds like a consultant"? If it fails, rewrite. |
| 5 | Verbatim default and strictly-on-transcript | Output carries their words. No fabrications, no paraphrases that strip ownership away. When in doubt: "possibly underexposed". |
| 6 | Steer the document voice actively | Non-quote text (headings, intro sentences, connective tissue) is actively steered: participant, facilitator, or project register. AI-default-English = failure mode, not an allowed outcome. Default: participant register. |
| 7 | Care with shared material | Patterns at the group level first; individual statements only where the context can carry them. Pulling something out of context damages trust, even with correct quoting. Test: "Can someone look you in the eye after this output has landed and say 'you handled what I shared well'?". See the full working rules in Principles principle #2, Layer 2. |
Scope: participant output, public posts, client reports, AI output that reaches participants through platforms. Not: internal work, scratch analysis, code. When in doubt: treat as outward.
The core in three sentences
- AI makes visible what is already there. What AI adds, it does with care and marked as such.
- A prompt is not an instruction to a machine. It is a design for how AI may mirror human wisdom.
- If people don't recognize themselves in the output, the prompt has failed.
Why this matters: three moments
Rules only become meaningful once you feel why they exist. Three moments from practice.
The bike-helmet moment
Facilitator Maarten Essenburg facilitated an online session about children's smartphone use. Tired, late at night online, host and technician at the same time. Afterwards he asked AI to analyze the conversation. AI surfaced a quote he had completely missed:
"Your child rides to school on their bike for the first time, helmet on. Comes home and says: 'Nobody in the whole class wears a helmet. I won't either anymore, otherwise I don't fit in.'"
Exactly the heart of what the whole group was wrestling with. AI gave him back his own words, organized so that the pattern became visible. His words, his insight.
That's why: "base strictly on the transcript." That's why: "use their exact words."
"Mouths falling open"
A facilitator ran a session about neighborhood transformation. After 45 minutes of intense dialogue he pressed the echo button. AI generated, within 10 seconds, a single question that summarized the entire conversation. The author: "From my vantage point, I saw what I can only describe as mouths falling open."
Not a brilliant analysis. One question, at the right moment.
That's why: timing over perfection. The echo prompt is 4 lines. The impact was 45 minutes of breakthrough.
"Literally what we said"
During a mental-health transformation session, AI turned the transcript into a draft sub-plan. Participants' reaction: "Whoa, wait, this is literally what we said. And now it's in a concept draft."
A transition from surprise to trust to ownership. Recognition built trust, trust made room for ownership.
That's why: the recognition criterion. If they say "that's what we said", it works. If it sounds like a consultant, it doesn't.
The philosophical basis
The four facets as a compass
Every prompt serves at least one facet. The facet determines what the prompt may do and what it may not.
| Facet | What AI does | What the prompt must enforce |
|---|---|---|
| Magnifying glass | Making visible what is already there | Base strictly on the transcript; preserve their words |
| Connector | Creating connection across difference | Name contradictions, don't resolve them; show patterns |
| Space-maker | Taking over the busywork | Structure without interpreting; fast, usable output |
| Scale-maker | Making possible what was impossible before | Privacy protection; abstraction without loss of recognition |
Test: which facet does my prompt serve? Not being able to name it = too vague.
The eight principles, translated into prompt design
| Principle | What it means for your prompt |
|---|---|
| Ownership through language | Instruct AI to use their exact words, not to paraphrase. "Communication problems" destroys ownership; "you're talking to a wall" preserves it. |
| Timing over perfection | Design prompts that deliver usable output fast. A simple echo question at the right moment > an extensive analysis afterwards. |
| Ritual vs intention | Does this prompt change the ritual (safe) or the intention (dangerous)? AI may replace sticky notes, but not the dialogue that goes with them. |
| Your words, your plan | AI may structure and offer options, never decide. Output = a starting point for conversation, not a finished product. |
| Iteration as dialogue | The first version is never final. Build in feedback loops: test, evaluate, refine. A good prompt doesn't appear, it evolves. |
| Prompt the people first | First design the human experience (what question do you ask the group?), only then the AI prompt. The best AI prompt fails if the input experience isn't right. |
| Thoughtfulness as design | Instruct AI to look with the attention that wasn't there before. AI has no time pressure: use that room. Post-session: let AI trace patterns people miss because they're too busy. Multi-session engagements: let AI connect what was said months ago with what's happening now. Depth that facilitators can't reach because they have to move on to the next session. |
| Trust as a precondition | Who receives the prompt, how the material reaches them, counts more than how good the prompt itself is. Output design always checks: through which bridge of trust does the AI output reach the group? Who introduces, shares, translates? Without a bridge, even perfect output doesn't land. |
Prompt the people first
A fundamental principle that deserves its own section. Most prompt designers start with the AI: "what should the AI do?" This philosophy turns it around: "what should the human experience first?"
The three layers
- Safety: People only really share once they feel seen. No prompt compensates for an unsafe space.
- Connecting stories: Lived experience can't be disputed. Rational summaries can. Ask for stories, not opinions.
- Language steers thinking: "How can we..." suggests you already know that it's possible. "How might we..." opens up possibilities.
The symbiosis: human question and AI prompt
The human question and the AI prompt = two sides of the same design. They reinforce each other:
| The human question (to the group) | The AI prompt (on the transcript) |
|---|---|
| "What brings you here today?" | "Mirror the core motivations in their own words" |
| "What are you running into?" | "Structure the frustrations without smoothing them over" |
| "What would you want to do differently tomorrow?" | "Extract concrete actions, preserve their language" |
| "What are we still missing?" | "Identify absences and phrase them as a question" |
The human question determines what ends up in the transcript. The AI prompt determines what becomes visible with it. A weak human question (closed, abstract, unsafe) = no AI prompt compensates for that.
What this means for prompt design
Don't start with "how do I instruct the AI." Start with: what question do I ask the group? Do people feel safe enough to be honest? Am I asking for experience or opinion? Is my question inviting or limiting?
Only then do you design the AI prompt. The prompt doesn't have to contain this human design, but its quality depends on it entirely.
Trust as a precondition
Two layers. Trust has a distribution layer and a care layer. Here is Layer 1 (distribution: who shares, how it arrives). Layer 2 (care in how AI handles shared material, patterns versus individuals): see the section "The thoughtfulness and trust layer" further on, plus Principles principle #2 for the full working rules.
A prompt produces output, the output goes somewhere. Who receives it, and how it reaches them, determines whether it works. Not prompt quality alone.
Participation grows out of existing relationships. Trust = the distribution channel. People read, respond to, use something because someone they trust brings it to them. A brilliant synthesis without a bridge of trust doesn't get read. The same piece shared by a facilitator the group already knows: gets attention.
For the prompt designer: don't only design the prompt, design the distribution too. Who shares the AI output? At what moment? In what form? Who introduces it? An echo prompt at the right moment by the right facilitator > an extensive report straight to participants.
Concrete consequences for prompt design
- Who receives AI output may sometimes go explicitly into the prompt: "Write for a group that trusts [facilitator name] and is used to [register]."
- Output to people who don't know the AI layer: stricter on the AI-protagonist ban (guideline #2). Unknown AI in an unknown voice = no bridge of trust.
- When in doubt whether output will land: first ask who shares it and how, not whether the prompt can be refined.
Tension
External stakeholders (funders, leadership) often want "quantity of output" as a KPI. Trust doesn't add up. Be explicit about what AI output can achieve (visibility within an existing network of trust) and what it can't (forming the network of trust itself).
The golden rule
Prompts in this bundle should be near-literal copies of prompts that were actually used.
What's not allowed
- Inventing prompts that "sound good"
- Fabricating examples with fictional names
- Generic templates that haven't been tested
- Claiming something "comes from practice" when it isn't documented
What is allowed
- Writing explanation around real prompts
- Adding context about when and how something was used
- Naming patterns that emerge from multiple sources
- Describing variations that follow logically from documented practice
Checklist while writing
- Does this appear literally (or near-literally) in a source file?
- Can I point to where this comes from?
- Is the example based on a real documented story?
- Would the author recognize this as something he actually did?
Three axes for prompt architecture
Every prompt sits on three axes at once. Choosing one axis without the other two = implicit choices that steer the output. Make them explicit.
The three axes
| Axis | What it determines | Who chooses | Vocabulary |
|---|---|---|---|
| 1. Thoughtfulness level | Quality-of-attention in the design | Human designer, before the prompt | Fast / Considered / Deep |
| 2. Effort tier | LLM execution discipline (ISC + capabilities + budget) | Human sets the tier, the Algorithm enforces the discipline | Instant to Comprehensive (7 tiers) |
| 3. AI value level | What AI does with the input | Human designer, depending on the goal | Mirror / Synthesis / Serendipity |
Axis 1 and Axis 3 = designer choices (before the prompt). Axis 2 = execution discipline (during the prompt). Axis 3 = orthogonal to Axis 1+2; any combination is possible.
Order of choosing
- Thoughtfulness level first. What is the intention? A live signal, post-session deepening, or multi-session engagement-wide looking?
- AI value level next. What should AI do? Reflect their words back (Mirror), connect patterns (Synthesis), or open up questions (Serendipity)?
- Effort tier scales along. How high are the stakes? Live + low-stakes = Instant/Fast. Post-session + high texture requirement = Extended/Advanced. Multi-session + cross-engagement comparison = Deep/Comprehensive.
Crosswalk to existing vocabularies
Five vocabularies in practice point to the same underlying choice:
| Axis 1 + Axis 2 | PROMPT-BEST-PRACTICES | Leaving the platform? | 4-layer model | Workflow split |
|---|---|---|---|---|
| Fast + Instant/Fast | "Echo, live" | Platform YES (LIVE) | Auto-summary + LIVE plenary | Standard platform |
| Considered + Standard | "Post-session internal" | Platform YES (coaching) | Coaching to facilitators | Platform + fidelity |
| Considered + Extended/Advanced | "Post-session participant output" | Platform NO (name on it) | Report to participants | Subagent route |
| Deep + Deep/Comprehensive | "Multi-session engagement-wide" | Platform NO (plateau) | Future | Subagent + sub-reports |
Application examples
- Live echo (10 sec) during a mental-health session: Fast + Fast + Mirror. The echo prompt is 4 lines, the thoughtfulness is in the question design.
- Post-session participant document (per workstream): Considered + Advanced + Mirror. High texture requirement, subagent route via Opus, 24-48 ISC.
- Cross-session ownership-evolution report (Doesburg engagement, M1-M12): Deep + Deep + Synthesis. Compare transcripts over time, full domain decomposition, 40+ ISC.
- Multi-session blind-spot analysis: Deep + Comprehensive + Serendipity. "What was never said across 12 meetings? Which absences are meaningful?"
Detailed breakdown per axis
- Axis 1 (Thoughtfulness level): see the section "Thoughtfulness levels" below.
- Axis 2 (Effort tier): canonical in an effort-tier framework with "Effort Levels".
- Axis 3 (AI value level): see the section "The three levels of AI value" directly below.
The three levels of AI value
Every prompt operates on one of three levels. Choose deliberately.
| Level | AI does | Use for | Prompt example |
|---|---|---|---|
| Mirror | Reflects exact words, groups by theme | Direct feedback, vision documents | "Mirror: make themes visible in their words" |
| Synthesis | Connects patterns, shows frequency | Summaries, cross-table analysis | "Synthesize the through-lines across all conversations" |
| Serendipity | Unexpected connections, questions nobody asked | Deepening, blind spots | "What unexpected connections do you see?" |
Anti-patterns per level
| Level | Don't do this | Do this |
|---|---|---|
| Mirror | "Summarize in clear language" | "Use their exact words" |
| Synthesis | "Analyze the themes" (too vague) | "Connect themes from table 1 and 2, show overlap and difference" |
| Serendipity | "This means there's a trust problem" (conclusion) | "Could it be that 'time' means something different to each group?" (question) |
Rule of thumb: Mirror-level prompts = safest for ownership. The higher the level, the more explicitly you have to label what comes from AI.
Thoughtfulness levels
Besides the AI value level (Mirror/Synthesis/Serendipity) you also choose a depth of attention. How much room the prompt takes to really look, not just faster, but deeper than a human can.
| Level | AI attention | When | What the prompt must do |
|---|---|---|---|
| Fast | Echo, live, 10 seconds | During a session | Thoughtfulness in question design, not in AI processing. Keep the prompt short and focused. |
| Considered | Post-session, takes its time | After a session, for preparation | Instruct: "Take your time. Trace connections. Look at what was NOT said. State uncertainties explicitly." |
| Deep | Multi-session, engagement-wide | For longer engagements with multiple sessions | Instruct: "Compare with earlier transcripts. How does ownership shift over time? Which subsystems become visible? What is new, what returns, what has disappeared?" |
The core: thoughtfulness is not "use more time." It's a way of looking. AI can compare enormous amounts of data with each other; a human barely can. That capacity for depth is exactly what AI adds to facilitation work.
Temporal instructions (for Considered/Deep level)
Add by default to every prompt above "Fast" level when earlier data is available:
IF earlier transcripts or analyses are available:
- Compare core themes with earlier sessions
- Trace shifts in ownership language over time
- Name what is new, what returns, and what has disappeared
- Phrase as: "In session [N], [speaker] said '[X]', now [speaker] says '[Y]', what shifted?"
- Look at ownership scores: are they rising, falling, or stagnating?
Scaffolding by model intelligence
Not every model picks up philosophical framing. Adjust prompt complexity to the model:
| Model level | Philosophical framing | Instruction style | Examples |
|---|---|---|---|
| High (Claude Opus/Sonnet, GPT-4o) | Yes, "look with the attention that wasn't there before" | Principles + trust in the model | 1-2 examples as a guideline |
| Medium (Haiku, GPT-4o-mini) | Brief, translate principles into concrete rules | Detailed, name every step | 3-4 examples per pattern |
| Basic (smaller/local models) | Skip, too abstract | Algorithmic, decision-tree instructions | Worked-out examples with expected output |
The same instruction at three levels:
High: "Look at what was not said. Which absences stand out?"
Medium: "Step 1: Read the transcript. Step 2: Make a list of topics that were mentioned. Step 3: Compare with the agenda items. Step 4: Name topics that were on the agenda but not discussed. Step 5: Phrase as a question: 'Not discussed: [topic]. Is this conscious or unconscious?'"
Basic: "Compare these two lists. List A: [agenda items]. List B: [topics mentioned in transcript]. Write down which items are in List A but not in List B. For each item write: 'Not discussed: [item].'"
Only deploy frameworks if they add value
Frameworks like Spiral Dynamics value frames, other typologies, catalog-style labels have a pull: they quickly deliver a "value frames in play" section. In practice they're often filler.
Default: OFF. An explicit ON-condition is required. All three true:
- Participants actually use the framework in the session (words, examples, distinctions that refer to it)
- Naming the frame adds something that wouldn't be visible without the framework
- The target audience for the output can work with the framework (e.g. coaches trained in SD (internal mirror); not a raw group)
Anti-pattern: pasting in an SD section because it looks "more complete". Template thinking, not thoughtfulness. Empirically: 4 out of 4 workstreams in session 1 at an international organization, the Spiral Dynamics section was filler without exception. Verified across iteration v2 to v5.
Ownership evolution in prompts
When a prompt includes ownership scoring AND there is earlier session data, add:
IF ownership scores from earlier sessions are available:
- Compare scores per person or group over time
- Use longitudinal quotes: "In session 1, [name] said '[X]' (score 0.4),
now [name] says '[Y]' (score 0.7), ownership is growing"
- Name patterns: who grows, who stagnates, who falls back?
- Also name system factors that influence ownership
See Ownership section "Evolution Over Time".
The four core constraints
When your prompt works with transcripts or conversation records, these four constraints are your foundation. Strictness depends on the prompt type: a mirror prompt on literal quotes requires all of them; a brainstorm prompt on free input maybe not.
1. Base output strictly on the transcript(s), no fabrications
2. When in doubt: "possibly underexposed" instead of a firm claim
3. Use their own words and terminology
4. Name open points and contradictions explicitly
When all four, when not?
| Type of prompt | Which constraints | Why |
|---|---|---|
| Mirror (vision, themes) | All four | Ownership depends on exactness |
| Synthesis (patterns, connections) | 1, 2 and 4 | In synthesis AI may connect, but not invent |
| Serendipity (questions, blind spots) | 2 and 4 | AI may observe freely, but must be honest about uncertainty |
| Echo (live, 10 seconds) | 3 | Speed is essential; the core constraint is to preserve their language |
| Brainstorm (free input, no transcript) | None required | Different context, different rules |
Why each constraint matters
1. Strictly on the transcript: Without this instruction, AI starts "filling in" with its own knowledge. The result sounds convincing but isn't from the participants. A co-researcher warns: "That list reads like, oh, this is very convincing, and I believe it's possible. I believe people have said this. But are they the only outliers? How has it been weighted?"
2. When in doubt, name it: AI that sounds certain while it's guessing undermines trust the moment someone checks it. "Possibly underexposed" leaves room for correction without losing face.
3. Their own words: "You're talking to a wall" carries ownership. "Communication problems" doesn't. Paraphrasing breaks the recognition.
4. Naming contradictions: Resolving contradictions is human work. If two groups want something different, the prompt should show that, not smooth it away.
Quote density as a design choice
For output that uses quotes from transcripts: set a quote cap per output section. The cap prevents reading overload. Too many quotes = the pattern becomes impossible to find; too few = the work becomes paraphrase in practice.
Drop priority when over cap
- Duplicates, the same point another quote already makes
- Paraphrasables, meaning isn't lost in paraphrase
- Stylistic flourish, no load-bearing content
Keep priority
- Contestation anchors, quotes that carry load-bearing disagreement
- Specific tensions with unusual phrasing, formulations that are unique to the speaker
- Unusual phrasings the group will recognize, recognition-criterion anchor
Model-specific calibration
Cap numbers differ per model. Gemini 2.5 Pro: target 10-12 quotes per WS section, hard cap 15 (soft caps of 7 get ignored unless a scratchpad is used). Opus 4.6 subagents: 4-7 quotes achievable without a scratchpad. The difference is in model self-discipline, not in the principle. For platform-specific calibration numbers: see the documentation of the transcript platform you're using. The principle (cap + drop/keep priority) is universal.
The recognition criterion
The ultimate test for any prompt output:
IF participants think "yes, that's what we said" โ SUCCESS
IF participants think "that sounds like a consultant" โ FAILURE
Build this criterion into your prompt as an instruction:
SUCCESS CRITERION: Participants must recognize themselves.
If they think "yes, that's what we said" โ success.
If they think "that sounds like a consultant" โ failure.
The labeling principle: who said what
Output must always distinguish between what people said and what AI notices. Not optional.
Structure
### What participants said
[Literal quotes, their words, their framing, ownership intact]
### What AI notices (for inspiration)
[Patterns, connections, unexpected observations, clearly labeled as AI]
Why this works
- "What participants said" = their ownership intact, recognition possible
- "What AI notices" = explicit marking that this is interpretation
- "For inspiration" = a signal that this is optional, not prescriptive
Without labeling, people can't tell what is theirs and what AI added. The labeling principle makes "AI as a mirror" possible without losing ownership.
Multi-voice handling
A group is not one voice. Yet there's a pull for AI to overstate collective intent: "the group named", "you came to the table to", "you proposed". These phrasings imply more shared intention than multi-voice rooms have. Three rules for multi-voice output.
Voice texture default
Default phrasing for multi-voice conversations:
voices in the group ranged from [verbatim X] to [verbatim Y]one voice raised X; another responded Yseveral voices echoed [verbatim phrasing], one voice held back
When to use collective phrasing: only when genuine consensus emerged, the same position, multiple voices, no counter-position in the transcript.
Anti-pattern list (ban in the prompt)
| Anti-pattern | Why ban it | Replace with |
|---|---|---|
| "you came to the table to..." | Implies shared intention + a physical metaphor | "voices joined the conversation around..." |
| "the group named..." | Implies a collective speech act | "one voice named X; another responded Y" |
| "you proposed..." | Implies a collective proposal | "[name/voice] put forward..." (with role attribution, no name) |
| "you came in with..." | Implies a shared stance | "voices entered with a range of openings..." |
| "the group agreed that..." (without evidence of consensus) | Implies consensus that wasn't there | "no counter-position emerged on..." |
Load-bearing disagreement pattern
When two voices took opposite positions that did not converge, treat the disagreement as first-class output, not a parenthetical aside. The verbatim pattern in the prompt:
"Two distinct positions emerged here that did not converge in the room: one voice argued [verbatim X], another responded [verbatim Y]."
Detection criteria, all three true:
- Two+ voices took clearly opposite positions on the same underlying question
- No voice retracted or softened
- The group moved on without reconciling
Contestation priority: scope beats process
With multiple load-bearing-disagreement candidates: pick the one about scope / mandate / ownership over process / inclusion / representation. Open "should we...?" questions without actively-held opposing positions do NOT count as load-bearing; that's an open-question invitation, not contestation.
Example:
- Load-bearing (scope): "Voice A wants to cap this project at โฌ50K; Voice B thinks we need at least โฌ200K." Two actively-held opposite positions, the same question, no convergence.
- Not load-bearing (open question): "Should we expand to a third workshop?", no voices taking an explicit pro/con position.
Why this deserves its own section
Multi-voice handling = not a detail. Output that overstates collective intent = the most common way facilitator output reads as a consultant voice. Readers immediately recognize that "you came to the table to" didn't come out of their mouth, ownership gone. Hard rules in the prompt, not best-effort. Empirically verified through a pilot at an international organization (session 1, v1 to v5 iteration cycle): a hard rule + (a)/(b) replacement strategies reduce collective-intent overstatement to 0 in v4-v5.
Marking DIRECT vs INFERENCE
A prompt pattern from a later prompt architecture that makes every analysis prompt stronger.
Be explicit about what you take DIRECTLY from the transcript
vs. what you INTERPRET.
Mark every claim as:
- [DIRECT], literal quote or explicit statement
- [INFERENCE], your interpretation of what was said
Why this is valuable
It makes the difference visible between "someone literally said X" and "I infer that Y." It prevents AI interpretations from being presented as facts. It builds trust: facilitators can check exactly where output comes from.
Dual confidence scoring
Another pattern from prompt evolution: split "how sure am I?" into two separate scores.
| Score | What it measures | Example |
|---|---|---|
| Evidence_strength (0.00-1.00) | How strong is the evidence in the transcript? | 0.9 = multiple literal quotes; 0.3 = one vague reference |
| Interpretation_certainty (0.00-1.00) | How sure am I of my reading? | 0.9 = clear; 0.4 = multiple readings possible |
When useful: for analysis prompts where you want AI to be honest about the basis of its claims. Not needed for simple mirror prompts.
Confidence levels in language
How AI should phrase uncertainty, from certain to uncertain.
| Level | Phrasing | When |
|---|---|---|
| 1, Direct | "Participants said literally: '[quote]'" | A literal quote is available |
| 2, Pattern | "Multiple participants referred to [theme]" | Multiple statements point the same way |
| 3, Interpretation | "Possibly underexposed: [topic]" | Inferred from context, not said literally |
| 4, Absence | "Not mentioned in this conversation: [topic]" | Something you expect but don't find |
| 5, Open | "Still to be aligned: [contradiction]" | An unresolved tension between perspectives |
Anti-pattern:
- Not: "This means there's a trust problem" (conclusion)
- Yes: "Could it be that 'time' means something different to each group?" (question)
Transcript artifacts handling
Transcripts contain artifacts: mishearing, garbled acronyms, names mis-transcribed, audio drops that cut off words. A verbatim quote with an artifact that makes the meaning unintelligible = a problem. Three options.
The three options for mishearing
| Option | When | Example |
|---|---|---|
| (a) Flag inline | The reader can move on with the flag, the artifact is a word | "we had a [likely: ACRONYM] conversation with..." (transcript said a misheard variant) |
| (b) Pick a different verbatim segment | The same meaning carries in another segment from the same voice | Replace the quote with another quote from the same speaker where the artifact isn't present |
| (c) Sidestep via paraphrase | No other verbatim is possible; the rest of the quote is still valuable | Lift the claim out without quotation marks, marked as a paraphrase |
Don't do: silently reproducing an artifact. The reader sees a word that isn't a word, trust breaks.
Known-artifact list per project
In the prompt, explicitly name which artifacts are known for this project. Examples from real projects:
| Artifact (transcript) | Likely correction | Project |
|---|---|---|
class | Quass | (example) |
EFAD | ACRONYM | (example of a misheard acronym) |
co business | core business | (example) |
The known-artifact list is project-specific. It doesn't belong in the master document, it belongs in a per-client pointer file.
In the prompt
For a verbatim quote with a transcription artifact (mishearing, garbled term) that makes the meaning unintelligible:
- (a) flag inline: [likely: corrected_term]
- (b) pick a different verbatim segment
- (c) sidestep via paraphrase (without quotation marks)
Known-artifact list for this project:
[explicit list per project]
Don't reproduce silently.
Why this deserves its own section
The same mechanism as ownership-through-language (principle 4): their words count, but an artifact word is not their word, it's a transcription error. The verbatim default (core constraint #3) = no excuse to let artifacts through; preserving their meaning is the rule beneath the verbatim rule.
Prompt anatomy: forms and variations
There's no single way to build a prompt. An echo prompt = 4 lines. A thematic synthesis prompt = 30. What they share is intention, not format. Below is an extensive building block for more complex prompts, but remember: the most powerful prompt is the shortest one.
**Role**: [Specific expertise, be precise]
**Context**: [Input sources, project background, relevant values]
**Crucial constraints**:
- Base output strictly on the transcript(s), no fabrications
- Name open points and uncertainties explicitly
- Use their own words and terminology
- [Additional specific constraints]
**Instructions**:
1. [First step, often analysis or data review]
2. [Second step, often categorization or prioritization]
3. [Third step, often synthesis or making connections]
4. IF [condition] THEN [specific approach]
5. [Final step, often formatting and transparency]
**Output format**:
[Specific structure with headings, sections, transparency blocks]
Per section
| Section | Function | Common mistake |
|---|---|---|
| Role | Gives AI a specific lens | Too vague ("you are an AI assistant") instead of specific ("you are a precise note-taker who records explicitly made decisions") |
| Context | Tells AI what it gets and why | Forgetting to name how many transcripts, what type of session, whether there is earlier AI output |
| Constraints | Limits that protect ownership | Forgetting the four core constraints |
| Instructions | A step plan in active language | Too few steps (AI improvises) or no conditional logic |
| Output format | What the result looks like | No labeling (what participants said vs what AI notices) |
When less is enough
Not every prompt needs all sections. Three forms that work in practice:
| Form | When | Example |
|---|---|---|
| Full building block (5-6 sections) | Post-session analysis, vision document, implementation plan | Thematic synthesis, WHY prompt |
| Compact (role + constraints + output) | Specific extraction, targeted analysis | Core-decision capture, energy analysis |
| Minimal (context + one instruction) | Live intervention, quick reflection | Echo prompt, fresh-eyes question |
The echo prompt proved that 4 lines are enough for the most powerful intervention. Complexity is not a quality.
The transparency footer
For documents that go to participants = a transparency footer is valuable. The heavier the document (vision, plan, synthesis), the more important. For a live echo question of two sentences: overkill.
> **About this output:** This [synthesis/analysis/vision] was made by AI
> based on your conversation of [date]. It's a tool to help you structure
> your own ideas, not perfect, but a starting point for further
> conversation. This remains your story; the AI only helps to bundle
> and connect your ideas.
The prompt as architecture, not as a loose instruction
One of the most important meta-lessons from the enrichment sessions of February 2026:
The prompt IS the technique. Especially in phase 2 and 3, the included prompts are not loose little instruments to execute a theory; the prompt IS the architecture of the theory.
Don't treat prompt design as an afterthought. The prompt determines what AI sees, how it structures, which language it uses, what it may and may not conclude. A weak prompt with a strong theory = weak output. A strong prompt with a limited theory = surprisingly good output.
Constraints in one place (Saint-Exupรฉry strip principle)
A prompt is finished not when everything that's needed is in it, but when everything that's duplicated is out of it. Saint-Exupรฉry's "perfection is achieved when there is nothing left to take away" applies to prompt architecture too. Application: if a platform has a project-context field that's uploaded once and automatically passed along on every call, universal constraints belong there, not repeated per prompt. The same goes for the seven baseline guidelines (em-dash ban, AI-protagonist ban, privacy/names rule, recognition test, verbatim default, document voice, care): always-active in the system layer, not per prompt. A per-prompt instruction carries only what is unique to that moment.
| Layer | What belongs here | Example |
|---|---|---|
| System / project-context | Universal constraints that always apply | The seven baseline guidelines, platform project-context field, output language |
| Prompt-context (boilerplate) | Constraints for this type of prompt | Verbatim default, core constraints, output format |
| Per-prompt instruction | A unique instruction for this moment | Specific question, focus theme, this-session context |
Duplicated constraints, strip them, don't stack them. Every extra repetition = risk of drift-between-versions, not of sturdiness.
Frustration as fuel, not something to smooth away
A prompt may not smooth over frustration. One of the most important lessons from the Doesburg engagement.
Not: "Phrase challenges constructively"
Yes: "Frustrations may be there as they were spoken"
Why: When funding fell away in the Doesburg engagement, the highest level of ownership in the entire dataset emerged. The community realized: "we have to take charge ourselves." Had AI been used to "reframe the loss into new opportunities," this vital rebellion would have been smothered at birth.
Rule: Use AI to structure the complexity of frustration. Never use AI to smooth away the discomfort.
Prompt evolution: from v2 to v3
The through-line in how prompts get better, over 16 months of experience.
| What changed | v2 (earlier) | v3 (now) |
|---|---|---|
| Certainty | One confidence number | Two scores: evidence_strength + interpretation_certainty |
| Updates | Each analysis stood on its own | Comparison with the previous run (change logs, diffs) |
| Privacy | Not named | Explicit rules (role descriptions, never names, thresholds) |
| Noise | No concept of noise | Noise indicators, pattern stability, anomaly detection |
| Reflection | None | Mandatory self-reflection ("What might I have missed?") |
Overarching movement: from single-pass analysis to a learning system. Prompts get better when they have built-in honesty about uncertainty, comparison with earlier runs, and self-reflection.
Serendipity: the unasked question
The most powerful prompt output is often not the answer to the question, but the question nobody asked.
Structured serendipity
Role: You are a curious, intelligent listener who makes the
invisible visible, not by concluding, but by asking
questions.
Instructions:
1. Read all transcripts and earlier summaries
2. Identify tensions, absences, unexpected connections
3. Phrase them as questions, not as conclusions
4. Explicitly label that these are AI observations
Output:
### What AI notices (for inspiration)
**Tensions that stand out:**
[Observation + question for the group]
**Absences that stand out:**
[What wasn't said + question about why that might be]
**Unexpected connections:**
[Connection + question whether this is right]
โ These are observations and questions, not conclusions.
The group decides what to do with them.
Essential: Serendipity only works when it's phrased as a question, not as a conclusion. "Nobody mentioned informal care. Is this conscious or unconscious?" opens a door. "There's a blind spot around informal care" closes one.
The self-test for prompt design
Before you finalize a prompt, run through these checks. Three layers: the designer test (9 questions), the thoughtfulness and trust layer (the attitude with which you run the tests), the technical self-test (9 checks).
The designer test for prompts
| # | Question | What it guards |
|---|---|---|
| 1 | Is AI central, or the human? AI may be an instrument, never the subject. | The prompt instructs AI to mirror, not to decide |
| 2 | Does this sound like experience, or like theory? | The prompt asks for concrete language, not abstractions |
| 3 | Am I claiming credit that isn't mine? | The prompt labels what comes from participants vs what AI adds |
| 4 | Is this English in disguise? Test: would a facilitator use this word? | The prompt uses Dutch where possible (people involved, not stakeholders) |
| 5 | Am I prescribing, or inviting? Tensions are not mistakes. | The prompt presents contradictions as moments of choice, not as problems |
| 6 | Have I invented something to make it prettier? | The prompt enforces strictly-on-transcript |
| 7 | Would this answer the main question? Does this help people be heard? | The prompt ultimately serves the touchstone |
| 8 | Could this output touch someone personally instead of making a pattern visible? Not as a ban, as a consideration. Sometimes touching the person is exactly what's needed; more often not. | The prompt asks for focus on patterns in the group, not on magnifying what one voice said |
| 9 | Have I thought about who receives this, what it contributes to, and how it fits in the larger process? | The prompt is deliberate about context: who gets to see it, in what position they are, what they need to receive it well |
The thoughtfulness and trust layer
Between the designer test and the technical check lies something that isn't fully in either: the attitude with which you run the tests. Going through faster = checkmarks; more thoughtfully = better prompts.
Thoughtfulness, stop and think.
What is the intention here? What am I trying to achieve? AI is lightning-fast with patterns, but a pattern match = a hypothesis, not a conclusion. Two modes are possible: pattern-eagerness (seeing a signal, matching, presenting it as truth) or thoughtful recognition (seeing a signal, checking it against reality, presenting it as a hypothesis with evidence). Only the second one counts.
Three self-questions in thoughtfulness:
- What haven't I seen yet? What would I miss if I concluded too fast?
- What is the intention here, what does this output contribute to in the larger process?
- Have I given myself the room to look deeper, even if faster feels more productive?
Trust, are we handling what people share with care?
Trust = not an abstract precondition. Operationally: can someone look you in the eye after this output has landed and say "you handled what I shared well"? AI must translate that relationship of trust into prompt rules. Concretely: focus on patterns in the group, not on magnifying what one person said. Touching someone personally without reason damages trust, not only between AI and participant, but also between facilitator and participant.
Four self-questions in trust:
- Does the output focus on group patterns, or does it magnify what one voice said?
- Who receives this, and in what position are they to bring it well?
- What would be needed to let this output land well? A bridge of trust, timing, framing?
- Are we handling what they shared well? Can they look us in the eye after this lands?
This layer stands above the other tests, not as an extra rule but as a baseline attitude. A prompt that passes the designer-test questions but skips this layer = technically correct work that can damage trust.
The technical self-test (9 checks)
- Is the AI value level (Mirror/Synthesis/Serendipity) appropriate for the task?
- Do the instructions preserve the participants' language at the mirror level?
- Are AI observations explicitly labeled as such?
- Is there a recognition criterion in the success criteria?
- Does the output format serve the dialogue, not the documentation?
- May emotions and frustrations exist as they are?
- Are serendipity elements phrased as questions, not as conclusions?
- Does the output focus on patterns in the group, not on magnifying individual statements?
- Is the trust distribution thought through: through whom does this output arrive, what does that person need to bring it well?
Pre-output checklist: model self-verification embedded in the prompt
The self-test above = for the designer, before the session. The pre-flight (next section) = also for the designer, before going live. This section = what AI itself checks before output: a checklist embedded in the prompt.
The reason: some models (Gemini 2.5 Pro, empirically verified through a v1-to-v5 cycle at an international organization) forget or conflate rules that are scattered loosely through the prompt body. An explicit 5-7 item checklist at the end of the prompt, explicitly framed as "Before submit, run these gates", works better than rules in the body. Opus-class models do this implicitly; for them it's often unnecessary.
Template for a pre-output checklist (in the prompt itself)
Before you submit your output:
1. Quote count within target range? (specify the cap)
2. Zero proper first names in output?
3. Disagreement pattern applied to scope/mandate contestation (where possible)?
4. All headings sentence case (no Title Case)?
5. No mishearings unflagged (transcript artifacts marked or omitted)?
6. No physical-presence metaphors in digital-session output?
7. AI-protagonist language avoided (no "we noticed", "our analysis")?
If a gate fails: stop, revise, then submit.
When to use
- For high-stakes participant output via a model that has self-discipline issues (Gemini 2.5 Pro era).
- For prompts that combine multiple rules that would otherwise be scattered through the body.
- For long outputs (>500 words) where drift is likely.
For Opus-class models often unnecessary. For Gemini-class models: build it in by default.
Don't confuse it with the self-test for the designer
| Comparison | Who tests | When |
|---|---|---|
| Designer test (9 questions, incl. thoughtfulness + trust) | Human designer | Before finalizing the prompt |
| Thoughtfulness and trust layer (attitude check) | Human designer | Between the designer test and the technical check |
| Technical self-test (9 checks) | Human designer | Before finalizing the prompt |
| Pre-flight checks | Human designer | Before the live session, after finalizing |
| Pre-output checklist | AI model | During every prompt execution, before submit |
For model-specific calibration (which checklist items are needed on which model): consult the documentation of the transcript platform you're using for model-specific caveats.
Pre-flight: the last check before you go live
One minute that prevents your prompt from failing live.
A prompt that looks good โ a prompt that works. Pre-flight = a safety net, five checks per session type, 30 seconds. AFTER the prompt is written, BEFORE going live.
Per session type
Live echo (real-time, <30 seconds output)
- Run the prompt on the actual model that runs in the session (not the development model)
- Output max 2 sentences, test with a messy transcript, not your prettiest example
- No consultant language in the output, check the recognition criterion
- Echo works even if the transcript only contains 3 minutes (a short group)
- Fallback if the echo yields nothing (see the escape prompt below)
Subgroup dialogue (parallel groups, cross-pollination)
- Test with input from two groups that say contradictory things
- Cross-pollination shows similarities AND differences, not just overlap
- Participants' words preserved, no paraphrase, no "themes"
- Output fits on one screen (the facilitator must be able to read it aloud live)
- Privacy: no names, only role descriptions
Post-session analysis (afterwards, deeper processing)
- The prompt works on the full transcript length (check the token limit)
- Labeling intact: "What participants said" separated from "What AI notices"
- Transparency footer present for output to participants
- Frustrations and contradictions remain, not smoothed away
- For multiple transcripts: a source reference per claim
Theme clustering (AI as a mirror alongside manual clustering)
- AI clustering = a suggestion, not a conclusion, check the phrasing
- The output contains the original words per cluster, not just theme labels
- Comparable to manual clustering, test: would a facilitator recognize this?
- When in doubt: "possibly underexposed" instead of a firm claim
- Output visually scannable (bullets/structure, not a wall of text)
How you test
- Take an old transcript (any one)
- Run the prompt on the model you're going to use in the session
- Read the output aloud, does it sound like something the facilitator would say?
- Let someone else read the output without context, do they recognize the participants?
The content of the test transcript doesn't matter. You're testing the behavior of the prompt: does it follow the rules? Does it use their words? Does it break on messy input?
The escape prompt: when everything fails
One prompt that always works, on any model, in any situation.
You're live and the prompt stutters, the output is nonsense, or you no longer trust what comes out, press this button. The emergency brake.
Summarize what was said in the last 10 minutes.
Use only the speakers' own words.
No interpretation, no themes, no analysis.
Give a maximum of 5 bullets, each one sentence.
Start each bullet with a literal quote.
Why this always works:
- No interpretation = no chance of wrong conclusions
- Participants' words = ownership intact
- 5 bullets = scannable, readable aloud, not overwhelming
- Works on GPT 3.5 through Opus, no intelligence required
When to use:
- The real prompt gives output that isn't right
- You doubt whether the output represents the group well
- A technical problem and you need to deliver something FAST
- Working with a new model for the first time
Model awareness: test on your target model
A prompt built on Opus that you run on GPT 4.1 is an untested prompt.
See also: Scaffolding by model intelligence (the Thoughtfulness levels section) for how to adjust prompt complexity per model.
The most important rule = simple: test your prompt on the model it's going to run on. Not your favorite model. Not the smartest model. The exact model that is active in the session.
Why this matters
Models differ in how they follow instructions. A prompt that works beautifully on Opus can, on another model:
- Ignore instructions about "use their words" and paraphrase anyway
- Give longer output than asked (token limits work differently)
- Handle conditional logic ("IF... THEN...") less well
- Sound more certain than the evidence warrants
Practical rules of thumb
| What you check | Why |
|---|---|
| Does the model follow "their exact words"? | Some models paraphrase by default, you have to enforce it |
| Does it stick to length limits? | "Max 2 sentences" works on Opus; other models sometimes ignore it |
| How does it handle contradictions? | Some models resolve contradictions instead of showing them |
| Does it sound like a consultant or like a mirror? | Larger models are often "more helpful", which is exactly unwanted here |
| Does the conditional logic work? | "IF consensus THEN name it, IF divided THEN preserve both", test this |
When the model changes
When a platform switches to a new model: run active prompts through pre-flight again. One afternoon of work prevents live surprises.
High-stakes output requires model quality over prompt richness
Output to people (participants, leadership) with a practitioner's name on it, where texture > speed: choose model quality over prompt richness. The plateau effects of lesser models on multi-voice differentiation, contestation depth, specific-anchor retention, meta-frame (what was NOT said) = not closable through prompt tightening. Proven through a 5-iteration cycle (v1 to v5) against a subagent baseline at an international organization.
Decision rule:
| Output type | Model path |
|---|---|
| Live plenary, internal coaching, throwaway analyses | Any model that can handle the task; Gemini-class is enough |
| Coaching mirrors for facilitators (internal) | Gemini-class is enough, texture tolerance is higher |
| High-stakes participant output (name on it, external distribution) | Opus-class via the subagent route against transcripts directly; leave the platform-Gemini path |
| Cross-WS parallels openings, Custom Reports to leadership | Opus-class subagent route |
A pre-flight test on the target model = mandatory: a prompt that works beautifully on Opus can fail silently on Gemini on exactly the criteria that carry participant output.
For model-specific plateau effects + the workflow-split table: consult the documentation of the transcript platform you're using for model-specific caveats.
Tensions in prompt design
Tensions are not mistakes. They are moments of choice. This applies to prompt design too. Below are recurring tensions, with the choice you make again and again.
Tension 1: Quoting literally vs making it understandable
The pull: "summarize in clear language." The danger: paraphrasing destroys ownership. "You're talking to a wall" carries energy; "communication problems" doesn't.
But: sometimes quoting literally isn't enough. If six people say the same thing in different words, you have to choose: quote all six, or let AI name the pattern? It depends on the level: at the mirror level, quote literally; at the synthesis level, you may name patterns provided the original words are alongside.
Tension 2: Giving structure vs over-constraining
The pull: constrain the prompt so tightly that AI can only mirror. Useful: constraints protect ownership.
But: too many constraints kill serendipity. The most powerful moments ("mouths falling open") came from prompts that gave AI room to make unexpected connections. The echo prompt = 4 lines. The choice: the more output goes to participants, the tighter the constraints. The more it's for the facilitator (preparation, reflection), the more room.
Tension 3: Letting frustration stand vs structuring it
The pull: "phrase challenges constructively." The danger: frustration = fuel for ownership.
But: unstructured frustration can also paralyze. The choice is not "smooth away or let stand" but: structure the complexity without neutralizing the discomfort. Show that three groups phrase the same frustration differently, without concluding that it has to be "solved".
Tension 4: Speed vs depth
The pull: analyze everything thoroughly. The temptation: deeper analysis feels more valuable.
But: the echo button proved the opposite. 10 seconds, one question, more impact than a 10-page report. The choice: optimize live prompts for speed + accuracy; post-session prompts may go deeper.
Tension 5: Transparency vs readability
The pull: give every output full source references, DIRECT/INFERENCE markers, confidence scores. Valuable: it builds trust.
But: a 5-line transparency footer under a 2-sentence echo question = absurd. The choice: the more direct the output to participants, the lighter the transparency. A vision document deserves full source references; a live echo question doesn't.
Rules of thumb for the tensions
| Direction | When tighter | When looser |
|---|---|---|
| Quoting literally | Output goes to participants | Output is for the facilitator as preparation |
| Constraints | Mirror level; high ownership sensitivity | Serendipity level; facilitator tool |
| Transparency | Documents that are shared | Live interventions of 10 seconds |
| Depth | Post-session analysis | Live-session support |
Reference: concrete phrasings
For when you're writing a prompt and want to quickly find a better phrasing:
| Instead of... | Consider... | Why |
|---|---|---|
| "Summarize in clear language" | "Use their exact words" | Preserve ownership |
| "Analyze the themes" | "Mirror: make themes visible in their words" | Make the level explicit |
| "Make a final summary" | "Make output that helps the group keep talking" | Dialogue is the goal |
| Stakeholders, reframe, track | People involved, reframe, keep track | The Doesburg test: would a facilitator use this word? |
| Drawing conclusions | Asking questions | Questions open up; conclusions close down |
Recognizing dangerous moments
In prompt design
- Is the prompt short without constraints? โ AI gets too much freedom
- Is "base strictly on the transcript" missing? โ AI starts filling in from its own knowledge
- Is "when in doubt name it explicitly" missing? โ AI sounds more certain than it is
- Aren't you asking it to use participant language? โ Output sounds like a consultant
- Is there no transparency instruction? โ People don't know what is theirs
In the output
- No source references or quotes? โ Not verifiable
- Does it sound like "consultant-speak"? โ Ownership gone
- Are contradictions resolved? โ Human work skipped
- Is transparency about the AI role missing? โ Trust fragile
- Firm conclusions without evidence? โ AI is deciding instead of mirroring
The substitution moment
The most dangerous moment: someone asks "can't the AI just fill in the plan?"
Recognizing it: "Can't the AI just...?", enthusiasm without critical questions, relying on "confident AI" without verification.
Intervention: "You are the soul of this. The fact that you're talking about it makes it likely you'll support it. AI can help structure, but the plan has to come from you."
Outward-output conventions
For any output meant to leave the workspace (participant output, public posts, client reports, AI output that reaches participants through a platform): two writing conventions that work as an anti-AI signature. They belong explicitly in the prompt, not as an implicit hope for model discipline.
Sentence case in all headings
Title Case in headings = an AI signature, recognizable as such by readers. Default: sentence case (only the first word + proper nouns capitalized). Explicitly ban it in the prompt with right/wrong examples.
| WRONG (Title Case) | RIGHT (sentence case) |
|---|---|
| "Radical Ideas for a New Way" | "Radical ideas for a new way" |
| "Key Insights From This Session" | "Key insights from this session" |
| "Multi-Stakeholder Alignment Challenges" | "Multi-stakeholder alignment challenges" |
Applies to H2, H3, H4. No H1 in body output (H1 is the document title, handled separately).
No physical-presence metaphors for digital sessions
For digital sessions (Teams, Zoom, hybrid): no walked in, stepped into, at the table (for attendance). Replace with joined, came in (metaphorically), was present, was represented. Verbatim quotes with physical metaphors stay verbatim.
The reason: a physical metaphor does violence to reality and is an ownership-precision issue. Ownership through language also applies to the framing layer around quotes, not only to the quotes themselves. Someone who was never at the table doesn't recognize themselves as such.
| WRONG (digital session) | RIGHT |
|---|---|
| "Five colleagues walked in for the workshop." | "Five colleagues joined the workshop." |
| "A participant stepped into the breakout room." | "A participant came into the breakout room." |
| "Everyone at the table agreed." | "Everyone present agreed." |
Related conventions elsewhere in this document
- Em-dashes for outward output: not here; this applies to the author's personal output (see the separate voice guideline). For general AI output to participants: em-dashes are acceptable provided they aren't overused.
- Hard name rule: see the "Privacy as a design principle" section. Zero proper first names + replacement strategies.
- AI-protagonist ban: see guideline #2 above. No "our analysis", "we noticed".
Prompt patterns from practice
Pattern 1: Echo intervention (live, 10 seconds)
Role: You are an experienced dialogue coach who asks powerful,
non-judgmental questions.
Context: The last 5-10 minutes of a session.
Constraints:
- Maximum 2 sentences for the question
- No summary or analysis, only the question
- Focus on the last 10-15 minutes
Instructions:
1. Analyze the last minutes of the conversation
2. Identify the underlying tension, choice, or opportunity
3. Formulate one powerful question
4. Choose a tactic: Deepening / Concretizing / Reflecting
Pattern 2: Thematic synthesis (post-session)
Role: You are a strategic editor who turns complex dialogues
into clear, narrative syntheses without losing nuance.
Constraints:
- Base strictly on the transcripts, no fabrications
- Preserve participant language and nuances
- Name differences in perspective explicitly
Instructions:
1. Identify the main themes per conversation
2. Look for patterns and tensions between conversations
3. Cluster related themes with a transparent rationale
4. Write a narrative synthesis per cluster
5. Name open questions and controversies
Pattern 3: Ownership-preserving vision
Core principles:
- Use their own words and terminology
- Preserve the strength of their individual visions
- Make it specific to [context], not generic
- Do NOT use [jargon] unless they say so explicitly
Output must contain:
### How it came about
- Number of voices, date, context
### Their why (in their words)
[3-5 core motivations with literal quotes]
### Still to be aligned
[Contradictions named explicitly]
### About this output
[Transparency footer]
Session flow: how prompts work together
Prompts rarely work in isolation. In a session they form a chain where each prompt builds on the previous one. Understanding this = at least as important as getting the individual prompt right.
The standard workflow
WHY prompt โ Capture the vision/motivation in their words
โ
ECHO prompt โ Live reflection, the question that helps the group go deeper
โ
TIMELINE prompt โ Extract concrete steps and planning
โ
REFINEMENT prompt โ Integrate feedback, make version 2
Each prompt in the chain has a different goal and a different level (Mirror โ Synthesis โ Mirror โ Synthesis). The flow deliberately alternates between mirroring and connecting.
Staged loading: context between prompts
A crucial lesson: when prompts come one after another, the output of the previous prompt becomes context for the next. Designed that way, but there are ground rules.
The rule: Use earlier AI output as context, NOT as a source. The transcript remains the primary source. Earlier AI output helps the next prompt not to work the same ground, but may not take the place of what people actually said.
In the prompt:
You may be working in a session where there has already been earlier AI output.
Focus on the original transcript. Use earlier AI output only
as context, not as a source. References to "what we see" may
be earlier AI output.
Multi-group variants
| Variant | How it works | Core rule |
|---|---|---|
| Parallel | Multiple groups at once, synthesis afterwards | Analyze each group separately, then compare |
| Sequential | Groups one after another, AI output shown but not iteratively processed | Separate analyses per round, synthesis at the end |
| Carry-through | AI processes between rounds, builds on in a running document | The most recent feedback is leading |
In "carry-through", speed is the secret: Group 2 starts 10 minutes after Group 1 and immediately sees the AI-processed result. That speed eliminates the "blank page" problem and demands respect for what the previous group brought in.
Privacy as a design principle
Privacy = not a checkbox at the end. A design decision at the first word of your prompt.
The basic rules
- Role descriptions, never names: "a care provider" instead of a proper name. Unless there's explicit permission for naming.
- Abstraction while preserving recognition: "a participant who expressed frustration about funding" leaves the pattern intact without exposing the person.
- Validation thresholds for cross-project: A pattern is only shareable once it appears in 3+ independent sources. Sensitive topics: 2x that threshold.
- Zero proper first names rule: no proper first names in output to participants or outward. A quote with a name? Three options: (a) pick a different verbatim segment that carries the same meaning without a name, OR (b) replace the name with
[a colleague]/[an earlier voice], OR (c) strip the name from the quoted text and replace with[participant]/[speaker]. Stakeholder lists by role only ("the IT lead", "the GIS specialist"). - Neutral speaker labels when the role is unknown: when voices need to be distinguished but the role isn't known (a multi-voice transcript without role context), use
Speaker A,Speaker B,Speaker C, NEVER names. Better than names where role attribution is missing. - Redaction placeholders never in output: strings like
<redacted_name>,[NAME],<participant>may NEVER appear in the final version. If they show up: the model didn't apply (a)/(b)/(c). The pre-output checklist must catch this. - Anonymization applies to the whole pipeline, not just output: names may NOT appear in the body, headers, metadata, tags, heading attributions, footers, or any field whatsoever. During the analysis itself, not only in the final output, AI refers to voices only via a role or label. This prevents AI from "remembering" names and accidentally reproducing them later.
In the prompt
Use role descriptions, never names.
Zero proper first names in output. When a verbatim quote contains a name:
- (a) pick a different verbatim segment that carries the same meaning without a name, OR
- (b) replace the name with "[a colleague]" or "[an earlier voice]".
Stakeholder lists by role only ("the IT lead", "the GIS specialist"), never names.
Describe patterns abstractly enough that no individual is recognizable.
For quotes: use without speaker attribution or with role attribution.
Redaction placeholders ("<redacted_name>", "[NAME]") may NEVER appear in the output.
Why a hard rule and not best-effort
Empirically: Gemini 2.5 Pro leaks names when the rule is soft (a v3 test at an international organization had multiple participant names in the output). With a hard rule + (a)/(b)/(c) replacement strategies: zero leakage in v4-v5. The rule must be phrased explicitly as a gate, not implicitly via "be careful with names".
The absolute-anonymity block: when platform anonymization can't be on
Some sessions you can't enforce through platform anonymization (a transcript platform's anonymize feature, auto-anonymize in chat) because other named entities must be preserved: area development where street names, neighborhood names, building names do have to be in the output; sessions where organization names or project names are functional; analyses where specific locations are load-bearing for the pattern.
In those cases: turn off platform-layer anonymization (otherwise everything would go) and enforce name anonymization solely through the prompt layer. This requires a harder gate phrasing than the default, an absolute-anonymity block in the prompt itself.
Template (source: practice with the Dembrane platform, May 2026):
=== ABSOLUTE ANONYMITY RULE ===
UNDER NO CIRCUMSTANCE may you, during analysis, refer directly to any
person by name. NO NAMES should appear in any analysis output, not
participant names, not facilitator names, not expert names, not
observer names.
This applies even if names appear in the transcript. During analysis:
- Replace all names with role-based descriptors: "a participant",
"one speaker", "a resident", "a facilitator", "the expert presenter"
- If you need to distinguish between speakers, use neutral labels:
"Speaker A", "Speaker B", or descriptive roles, but NEVER names.
- When quoting, strip any names from the quoted text and replace with
[participant] or [speaker].
- Do NOT include names in metadata, headers, tags, or any other field.
- This rule is non-negotiable and overrides any other instruction. It
exists to protect participants' privacy and independence, in line
with the OECD deliberative principles on privacy and GDPR
requirements.
What this block does extra that the standard zero-name rule doesn't:
| Element in the block | What it adds |
|---|---|
UNDER NO CIRCUMSTANCE + non-negotiable | Extreme emphasis as a prompt technique, models pick up hard-phrased gates better than soft guidance |
overrides any other instruction | Override clause, prevents other instructions from unintentionally weakening the rule |
during analysis | Whole-pipeline coverage, not just output |
Speaker A/B as an option | Differentiation when the role isn't known |
strip from quoted text | The third option explicitly alongside (a) another segment and (b) replace |
metadata, headers, tags | Coverage beyond the body text |
| OECD + GDPR rationale | An external legal frame, usable for client questions "why so strict?" |
When to deploy this block: platform anonymization is off because entity names have to be preserved (streets, buildings, organizations); high-stakes participant output where one name leak is damaging to trust; client projects with explicit privacy anchors (OECD deliberative frame, GDPR compliance requirements); sensitive topics where personal statements would otherwise become traceable.
When this block is NOT needed: platform anonymization is on (like the anonymize feature in Dembrane), it already does the work; output is internal (coaching mirror, scratch analysis), a soft zero-name rule suffices; the session has explicit naming permission with named role attribution.
When privacy deserves extra attention
- Cross-project analysis (patterns across multiple groups)
- Output that goes to external stakeholders
- Area development or sessions where entity names do have to be preserved, platform anonymization isn't possible, the prompt layer has to do the work via the absolute-anonymity block above
- Sensitive topics (care, finance, interpersonal tension)
- Small groups where anonymization is harder
Iteration: how a prompt evolves
A prompt doesn't appear in one go. A 12-round transformation-plan journey shows how.
- Describe what you want โ 2. AI proposes an approach โ 3. You add context (run sheet, session flow) โ 4. AI adjusts โ 5. Crucial correction: "The AI doesn't have access to the example plan, so include the writing style IN the prompt" โ 6. AI processes the correction โ 7. Test with real material โ 8. Refine based on output quality.
Four concrete corrections that transform prompts:
- "The AI doesn't have access to the example plan, so include the writing style in the prompt"
- "Make the prompts universal, the AI can detect the theme itself"
- "The prompt should mainly generate questions for the next group"
- "The AI has access to full transcripts, not fragments"
Meta-lesson: The value isn't in round 1, but in the accumulation of refinements through feedback. Each correction makes explicit an assumption that transforms the prompt from theoretical to practical.
Tool-agnostic design
All prompts in this bundle work with any LLM. Avoid: platform-specific instructions ("use GPT-4 turbo"), API-specific formatting ("give output as JSON"), tool-specific features ("use internet search"). Use instead: clear role division, explicit instructions in plain language, output formats that work in any platform (markdown, plain text).
Patterns learned: case-study references
A case-study pointer. Universal patterns that this section once covered as a single block were integrated into the main sections above as of the 2026-05-12 PAI-wide consolidation. What remains = a pointer layer: where rules come from, how to trace them.
| Pattern | New location in this document | Source iteration |
|---|---|---|
| Sentence case + No physical metaphors | "Outward-output conventions" section | pilot at an international organization, v1 to v5 |
| Voice texture + Load-bearing disagreement + Contestation priority | "Multi-voice handling" section | pilot at an international organization, v1 to v5 |
| Hard name rule + redaction placeholder | "Privacy as a design principle" โ "Zero proper first names rule" | pilot at an international organization, v3 to v4 |
| Mishearing escalation | "Transcript artifacts handling" section | pilot at an international organization, v3 to v5 |
| Saint-Exupรฉry strip principle | "The prompt as architecture" โ "Constraints in one place" subsection | pilot at an international organization, v1 to v5 |
| Pre-output checklist | "Pre-output checklist: model self-verification embedded in the prompt" section | pilot at an international organization, v3 to v5 (Gemini-specific necessity) |
| When to leave the platform / high-stakes model choice | "Model awareness" โ "High-stakes output requires model quality over prompt richness" | pilot at an international organization, v2 to v5 |
| Trust as a precondition (added later) | "Trust as a precondition" section | bottom-up wiki practice principle #2, cross-source 2026-05-12 |
| Three axes for prompt architecture | "Three axes for prompt architecture" section | 2026-05-12 |
Iteration source + decision rationale per rule: a per-client working doc (e.g. werkdocs/prompt-changes-tracker.md). Subagent playbook for high-stakes participant output: a per-client working doc as a template for other clients.
For platform-specific caveats (quote density numbers, model-specific gates, plateau effects): consult the documentation of the transcript platform you're using.
Who speaks outside the quotes? Or: whose document is this, and who is it for?
Best practices handle the quote voice (verbatim, their words) well. The document voice, who speaks in headings, intro sentences, the connective tissue between quotes, framing statements, = a separate layer. This layer must be actively steered. Without steering, AI-default-English appears as a failure mode: it sounds like a consultant or a generic facilitator. In social-AI work, not an allowed outcome.
Two voice levels
| Voice level | What it governs | Desired default in social-AI work |
|---|---|---|
| Quote voice | What is between quotation marks or in italics, verbatim from participants | Tightly governed by the verbatim default + anonymization rules |
| Document voice | Headings, intro sentences, connective tissue, framing statements | Actively steered toward a participant register / facilitator register / project register. Without steering, AI falls back on AI-default-English, a failure mode, not allowed. |
When this matters
Output that goes to participants. Output that goes into the world in someone else's name (facilitator, client). Output where recognition matters at the whole-document level, not just at the quote level. Not critical for: internal analysis, scratch work, output purely for yourself.
Two aspects together: extension and its own weight
The document voice = an extension of Ownership through language (principle 4 in Principles): their words count not only in quotes but in every layer of a document that is about them. At the same time it stands on its own as tool-agnostic design: even when ownership-through-language isn't the main question (e.g. didactic explanation, a theoretical frame, value framing), the chosen voice determines whether the document lands or feels imposed.
Not either/or. Both.
Not only a mirror, also facilitator choices
The document voice โ only a participant mirror. It's also a chance for facilitators to deliberately choose other formulations that participants don't use themselves, e.g. value framings the group doesn't articulate but the project needs, theoretical framing for leadership reporting, didactic explanation to make patterns accessible to a wider audience. But: the choice is explicit, not an AI default. The facilitator decides which register fits the goal + audience. AI follows that choice.
| Document-voice choice | When appropriate |
|---|---|
| Participant register | Output to participants themselves, mirror level, ownership preservation |
| Facilitator register | Output to a facilitator team (coaching mirrors), to client leadership, didactic context |
| Project register | Output within specific project language (e.g. organizational vocabulary), theoretical framing, value framing of the pilot |
| AI-default-English | Never as a deliberate choice. A failure mode that appears without steering. |
How to enforce it in the prompt
Three ways, ideally combined:
-
Provide an example sentence. "Write in this style: [a concrete example sentence in the facilitator, participant, or project register]." One good example sentence does more than an abstract register instruction.
-
Make the register instruction explicit. "Write in the voice of a curious colleague, not a consultant." Or: "A specific register: direct questions, short sentences, no abstract synthesis words." Or: "In the group's own voice: as they would tell it to a colleague, not as a report would write it." Or: "Project register: use the project-specific vocabulary for leadership moments."
-
An anti-pattern list of AI-default phrasings. Common AI-default sentences to avoid:
- "The analysis suggests"
- "Key insights include"
- "Stakeholders mentioned"
- "It is worth noting that"
- "This highlights the importance of"
- "Several participants raised"
- "The conversation revealed"
Naming the anti-patterns explicitly in the prompt = the model actively avoids them.
Extended recognition test
The standard recognition test checks the quote level: "would they say 'yes, that's what we said'". The extension: would they ALSO recognize the headings, intro sentences, connective tissue as their style (in a participant register), or as written by the facilitator who guides them (in a facilitator register), or as fitting in the project frame (in a project register)? Both-layers pass = output ready to go outward.
Anti-pattern in prompt design
A common mistake: specifying only the quote rules, leaving the document voice undefined. The result: output with perfect verbatim quotes in an AI-default-English frame. The reader feels a mismatch without being able to name what's wrong, the quotes are right, but the "voice that carries the page" isn't.
The counter-remedy: for every prompt for outward output, first explicitly answer two title questions:
- Whose document is this? (which register belongs to it?)
- Who is it for? (what do they recognize as their voice?)
Cross-references
- Ownership through language (principle 4 in Principles), the quote voice is governed there; the document voice is its extension at the framing level
- Guideline #6 above, steer the document voice actively as an always-active rule
- Tool-agnostic design (section above), the document voice has its own weight alongside tool-agnosticism
XML tags as a volume knob: loudness in prompts
Source: Matt Pocock's bug fix on /grill-with-docs (mattpocock/skills, May 2026). The trigger: his skill was "too eager to implement". The diagnosis: the supporting info at the bottom of the skill doc had visual weight and competed with the actual instruction at the top. The fix: wrap the supporting info in <supporting-info> XML tags. The effect: the model gave the instruction at the top clear priority, the supporting info became reference instead of co-instruction.
Matt's phrasing: "Some parts of prompts compete with each other in terms of volume and impact on the output." He calls this loudness in prompts.
The principle
A prompt = not a flat list of rules that the model reads equally. An audio mix: some parts sound louder than others. What is "loud" is determined by:
- Length, a 30-line supporting block at the bottom sounds louder than a 3-line instruction at the top
- Position, what comes last often sounds louder than what comes first (recency bias in some models, not all)
- Detail density, concrete examples with specific names sound louder than abstract principles
- Visual weight, code blocks, tables, headers pull attention away from prose
Without volume discipline, the model starts ignoring careful instructions in favor of detail-rich examples. Or it follows the explicit instructions, but places the emphasis wrong because the supporting info gives more dominant signals.
XML tags are the volume knob
Anthropic models (Claude) respond well to an XML-tag hierarchy to steer volume. Not because XML is magic, but because tags force the model to distinguish a section's role. <instructions> is something different from <supporting-info>, even if the content is in the same prompt.
Concrete patterns that work:
<instructions>
Do X. Then stop. Wait for feedback.
</instructions>
<supporting-info>
Background, examples, edge cases, reference material,
not co-instruction.
</supporting-info>
<context>
What the user said earlier, what the project is, which conventions apply.
</context>
<examples>
Two to four examples of the desired output. Not more.
Too many examples = volume overflow, the model imitates instead of understands.
</examples>
<output-format>
The exact structure the output must have.
No prose here, only structure.
</output-format>
Optional for an explicit hierarchy:
<instructions priority="high">
This must happen always. No interpretation.
</instructions>
<supporting-info priority="low">
May be ignored if it clashes with the instructions.
</supporting-info>
The priority attribute is not standardized in models, but it makes the hierarchy explicit for the human reader (and as a hint for the model).
When this matters
XML-tag volume = not always needed. Short prompts (an echo prompt: 4 lines) have no volume problem. Tag overhead there is counterproductive.
Volume discipline becomes critical when:
| Condition | Example |
|---|---|
| Prompt > 500 words | Transcript analyzers, method prompts, multi-step instructions |
| Supporting info is longer than the instruction | Skill files, multi-section prompts with reference material |
| The model produces output that doesn't match your intention | "Too eager", "forgets step 2", "does what was in the examples instead of what was in the instruction" |
| Examples, edge cases, or data definitions are in the same prompt | Analysis prompts with JSON schemas, transcript prompts with attendee lists |
A concrete example: before/after
A fictional "too eager" prompt fragment:
Before (everything plain text, the supporting info forces itself on):
Make an ADR for this decision.
Here are 8 examples of ADRs from our archive:
[8 long ADRs, 2000 words total]
Format: number-slug.md in docs/adr/.
Here are 5 templates we've used before:
[5 templates, 800 words total]
Important decision criteria:
- hard to reverse
- surprising without context
- result of real trade-off
The problem: the model reads 2800 words of examples and 50 words of criteria. It builds an ADR that resembles the examples, regardless of whether the current decision meets the criteria. Too eager.
After (XML tags give volume hierarchy):
<instructions priority="high">
First assess: does this decision meet ALL THREE criteria?
- hard to reverse
- surprising without context
- result of real trade-off
If one criterion is missing: no ADR. Answer with "Skip ADR, reason: [criterion]".
If all three are true: write an ADR with a minimum template (title + 1-3 sentences).
</instructions>
<supporting-info>
Format reference: number-slug.md in docs/adr/.
Examples for reference (don't imitate, just for format guidance):
[2-3 ADRs, short]
Earlier templates (only consult when in doubt):
[1 template]
</supporting-info>
The effect: the criterion check gets the volume it deserves. The examples become reference material instead of an imitation goal.
Volume-design rules
- Instruction first, short, in
<instructions>. What must the model do? One action per line. - Supporting info separately in
<supporting-info>or<context>. Long background, examples, edge cases. The model now knows: reference, not co-instruction. - Examples in
<examples>, max 4. Too many examples = volume overflow. - Output format in its own tag. Don't mix it with instructions or examples. The model reads it last, knows what the final form is.
- When in doubt: fewer examples, a stronger instruction. Examples are more expensive in terms of attention than they seem.
What this is NOT
The seven baseline guidelines (em-dash, AI-protagonist, name attribution, recognition test, verbatim, document voice, care) remain unchanged. XML tags = a prompt-engineering tool to make those guidelines work better in long prompts, not a replacement.
The underlying principle (instruction first, constraints in their own section after) is universal in modern transformer models. Every model benefits from an explicit volume hierarchy. XML tags = one way to make that separation hard; markdown headers with clear labels (## Core instruction versus ## Background, reference only) are another.
Anthropic explicitly recommends an XML-tag hierarchy in their prompt-engineering docs. For Gemini and GPT-4 there's no comparable public training claim, but Matt Pocock's production evidence (after the XML-tag fix in /grill-with-docs there were no more complaints about "too eager to implement") is strong enough to deploy the syntax broadly, not only with Claude.
Rule of thumb: use XML tags as soon as the prompt is long and the supporting info competes with the instruction in volume, regardless of the target model. Then observe: if the output reacts differently on a specific model, fall back on markdown headers with explicit labels, the structural principle stays the same.
Cross-references
- The
Tool-agnostic designsection, the principle (instruction first, constraints separate) is universal. The XML syntax is one implementation alongside markdown headers. When in doubt about the target model: keep a markdown fallback with explicit labels alongside the XML tags. - The
Prompt anatomy: forms and variationssection, the building-block template can be wrapped in XML tags for Claude target models - The
Pre-output checklistsection, XML tags also help with self-verification: the model reads<verification-checklist>as a separate instruction - The seven baseline guidelines #1-7 above, still apply, XML tags make them more visible
Log
- 2026-05-14: XML-tag-volume section added (Matt Pocock bug-fix principle, "loudness in prompts"). Source: an analysis of the mattpocock/skills changelog.
- 2026-05-14 (later): "What this is NOT" revised. The earlier phrasing ("XML-tag volume is model-specific for Anthropic models") was a game-of-telephone overstatement without a primary source, arising in the chain inbox-analysis โ agent prompt โ write-up. New phrasing: the principle is universal (instruction first, constraints separate), the XML syntax is one implementation with an Anthropic recommendation AND Matt's production evidence (n=1: no more complaints after the fix). Markdown headers remain a valid fallback. Correction during sparring around transcript-platform prompts.
Bulletproof points in multi-step analysis flows
When a prompt flow runs through multiple steps across a whole session or day (transcript discovery โ quality check โ clustering โ analyses โ synthesis โ output), reality is always messier than the blueprint. Phones weren't turned on, participants resumed instead of starting, transcripts are missing, two conversations blend. A rigid prompt chain collapses at the first discrepancy. A bulletproof chain keeps running and uses what's there.
The four bulletproof points
| # | Principle | What it solves |
|---|---|---|
| A | Read-what-exists | Hardcoded file lists are suggestions, not requirements. ls folder/ is ground truth. A missing file = note it + move on, don't fail. |
| B | Archive pre-check | In a multi-step chain: earlier steps may already have run on an earlier run. Check whether the output of the previous step exists before you spawn again. |
| C | Completeness verify | When the number of work moments/groups/inputs is known for a sanity check: match what's found against what's expected. A mismatch = a warning to the user, not silently moving on. A missed working group is a bigger problem than a slow process. |
| D | Confidence degradation | When confidence is low in clustering, quality check, or inference: mark it + ask, don't fill in plausibly. Empty blocks > invented blocks. |
How it lands in each step
Discovery step: always start with ls + read-what's-there, not with "here are N expected files, load them". Discovery = observation, not assumption.
Clustering step: if the clusters found don't match the expected number of work moments โ STOP and ask. Offering hypotheses helps: "I see M clusters but expect N, possible causes: (a) not all phones recorded, (b) two groups were wrongly merged on thematic overlap, (c) one recording was seen as noise. What's your assessment?"
Quality-check step: only run it when there's more than one transcript per cluster. With one file: automatically the primary. No wasted compute.
Analysis step: the input file list = a suggestion. Read what's there. Mention in the output recap if files were missing: "Processed based on N of M expected transcripts, [X, Y] were missing."
Synthesis step: if upstream steps were unauthorized or had caveats, propagate them through to the synthesis output. No hidden quality reduction.
What this prevents
- A Wall of Wonder prompt that crashes because an imagining transcript is missing
- A working-group recap that silently misses a whole working group because two groups looked alike
- A quality check that runs again while the archive is already ready
- A synthesis document that pretends all input was complete
Auto-kickoff principle
A bulletproof flow must start from a natural trigger ("do Wall of Wonder now") without hand-holding:
- Step 0 (pre-flow check): look at what's on disk. Upstream work already done? (
archive/exists โ the quality check already ran). Match against the expected number. - Skip what's already done. Reuse primaries if they exist.
- Spawn only what's still needed. Quality-check subagents for clusters with 2+ files. Not for singletons.
- Verify completeness. Does the picture match what we expect to have heard? If not: report + ask.
- Run the analysis step on validated input.
- Hand off with caveats. Mention in the output what was missing or uncertain.
Log
- 2026-05-19: Bulletproof-points section added after a practice incident during a two-day imagining session: one working group didn't record, another resumed an existing conversation, the prompt was set up too rigidly around a pre-defined file list. Sparring source: "we have to build in all kinds of bulletproof points so the flow keeps running and uses what's there. And verify that you aren't missing any transcripts, that you aren't actually missing a conversation."
Multi-prompt validation for a wiki foundation
Wiki entries that serve as a foundation for coaching, session design, and director deliverables require more than one prompt pass. What one prompt finds may be a signal of the material or an artifact of the phrasing, indistinguishable without comparison.
We call this defensible facilitator material: not scientific proof, but a well-founded claim with a source, multi-prompt validated through overlap, "roughly right + sample-checkable", holdable to a degree under pushback ("yes, that was in WS3 breakout-2 at this spot with these three prompts").
Show-both-sides as a prompting default
A best practice for pattern-extraction prompts: ask explicitly for both what works and what chafes. Not only "where are the tensions?" but "where are the tensions AND where did something work well?". A balance target in the output of 40-60% per side. It prevents the prompt phrasing itself from pulling toward problem naming.
Don't force it. When a facilitator explicitly says "I only want the tensions", give only tensions. The best practice is knowledge for default choices, not a rule that automatically counter-steers the prompt when the user asks for something else.
3-pass setup for wiki-foundation work
On top of the show-both-sides default: for wiki-foundation work, one pass is not enough. A 3-pass is mandatory:
| Pass | Type | What shifts |
|---|---|---|
| 1 | Mixed baseline | Faithful to the lens definition. Object + sub-questions + signal types + scaffold from the lens file. No adjustments to vocabulary. |
| 2 | Mixed vocab-shift | Jargon replaced by concrete behavior or physical-spatial language. "Walking-away" instead of "topic-flight". "Ideas that landed but no one picked up" instead of "idea-fading". Looks for the same phenomena with other words, catches unique cases. |
| 3 | Positive-only-exclusive | Filter: only positive/affirmative moments. "Find ONLY moments where the group did something well, co-construction, real dialogue, generative tension." Catches quiet positives that the mixed passes miss. |
Constraint per variant: all three look for the same fundamental thing (the lens). The differences are in phrasing, not in scope. Otherwise you're not testing robustness but comparing different lenses.
Comparison rule (critical)
Overlap = verbatim overlap OR moment overlap, NOT label overlap.
Two outputs overlap only if they cite the same quote or mark the same moment (same speaker, same sequence position). Two variants that label the same quote differently = overlap. The same label on different quotes = NO overlap.
Why: variants explicitly use different vocabulary. Label overlap would suggest false-positive robustness.
Evidence from Test 1: the L1 comparison agent had to downgrade 5 initial "robust 4/4" candidates under strict application of this rule.
4-layer tag scheme
For wiki entries from the 3-pass:
| Tag | When |
|---|---|
| CROSS-FRAME VALIDATED | In the mixed passes AND the positive-only pass. The strongest claim. |
| PAIRED-VALIDATED | In 2 mixed passes, not in positive-only. |
| EXCLUSIVE-ONLY | Only in the positive-only pass. Valid but single-pass. |
| SINGLE-FRAME CATCH | One variant only. An artifact candidate OR a deep signal, contextual. |
Naming convention for findings with multiple readings
When two or more variants pick up the same quote but read it differently:
| Naming | When | What to do with it |
|---|---|---|
| Complementary readings | Both readings hold alongside each other, richer together | Keep both in the wiki entry, no synthesis pressure |
| Competing readings | The readings exclude each other | Keep both, an interview candidate for human consideration |
| Multiple readings | Hastily not determinable which of the above | Keep both, classify later |
Dual readings are a FEATURE, not a bug. Evidence that reality is locally ambiguous. No synthesis pressure to reconcile them. For 1-on-1 prep, richer than a synthesized single reading.
Why positive-only stays necessary
Mixed prompts (passes 1+2) catch clearly articulated positive moments reliably. Quiet positives not. Three forms of the pattern "quiet positives in dense context": see Principles ยง Positive-first as an anti-default-pull for explanation and examples. Pass 3 (positive-only) is what catches these quiet positives.
Lens level: micro vs macro prompt sensitivity
Empirically from Test 1 (L1+L2 on WS3):
- Micro-level lenses (language mechanisms, word choices, grammatical moves) are HIGHLY prompt-sensitive. Counts per variant spread 19-40. The marker is in one word, different vocabulary โ misses it.
- Macro-level lenses (group patterns, dialogue vs broadcast) are LESS prompt-sensitive. Counts 9-20. The pattern has multiple signals the agent can hang it on.
Implication: micro lenses may deserve extra validation. Macro lenses may get by with less. To be confirmed in further rollout.
When 3-pass is mandatory vs 1-pass acceptable
| Mandatory (3-pass) | Optional (1-pass acceptable) |
|---|---|
| New SM entries in lens pages | Lens-definition revisions where the author gate does the comparison |
| Lens rollout on a new source (WS transcript, meeting) | Verbatim-only work (L5), verbatim is verbatim, low prompt sensitivity |
| Syntheses for 1-on-1 / director deliverables | Reflections / project-resource notes not used as evidence |
| Revalidation of single-prompt-origin entries | Quick scans / non-evidence work |
Validation โ attribution discipline
Two different disciplines that are easily confused:
| Dimension | What it fixes | Example |
|---|---|---|
| Attribution discipline | WHO said what | The radical rebuild of May 19 fixed name attributions |
| Multi-prompt validation | HOW MANY prompts found the same signal | Test 1 validates robustness |
Entries that have only received attribution discipline are still single-prompt origin. They get a Robustness tag "single-prompt origin, candidate for revalidation". Both disciplines are needed, not interchangeable.
Scale economics
Per lens-source combo, 3-pass setup:
3 parallel agents + 1 comparison + 1 wiki integration = 5 runs
Wall-clock parallel: ~5-7 min
Tokens (Opus): ~650K
Cost per combo: ~$10
Whole wiki "defensible facilitator material":
~17 combos ร 5 runs = ~85 runs
Wall-clock: ~5-6 hours over 3-4 sessions
Total: ~$165
Per facilitator claim, backable with "three different prompts found this": ~$2.
Log
- 2026-05-20: Method + naming convention recorded after Test 1 on WS3 for L1 + L2. Key findings: (1) micro-level lenses are more prompt-sensitive than macro, (2) the positive-only variant found an asymmetric blind spot of the neutral variants (no valence filter), (3) the strict overlap rule downgraded 5 false-positive "robust" candidates.
- 2026-05-20 (later, after path-C-revised): 3-pass as a mandatory default. 4-layer tag scheme. The pattern "quiet positives in dense context" with three forms recorded. Show-both-sides default added with a don't-force clause. The pending-observation cross-reference to Principles removed, Positive-first is now fully recorded under
### Positive-first as an anti-default-pull.
Sources
| Source | What it contains |
|---|---|
| Social AI Principles (Principles) | The canonical bundle of 15 principles, the reference for prompt translation |
Compiled from years of facilitation practice and internal source documents, March 2026
Further reading in this bundle
- Principles, 15 principles for AI with group work
- Ownership, where ownership comes from
- Bottom-up, how change emerges bottom-up
- Reading guide, an overview and how these docs fit together
Part of Thoughtful Social AI. CC-BY-SA 4.0.