---
name: ai-evaluation-sprint
description: "Plan a bounded 48-hour evidence sprint that turns an imminent AI decision — vendor choice, build approval, workflow commitment — into a recommendation memo with an explicit confidence level, by assigning one owner per risk area and keeping every evidence request tied to the decision. Use when a decision must be made within days, momentum is real, but the supporting evidence is scattered across demos, notes, and inboxes. Not for first-pass screening of a new AI idea (use ai-opportunity-triage) or for open-ended evaluation programs with no decision deadline."
---

# AI Evaluation Sprint

Plan and structure a 48-hour evidence sprint for an imminent AI decision. The artifacts
are a sprint plan — one named owner per risk area, bounded evidence requests, an hour
budget — and a recommendation memo skeleton the sprint lead completes at hour 48 with
the call, an explicit confidence level, and named open risks. The operating principle:
the sprint converts an open decision window into evidence; it does not extend the window.

This workflow is published by Sophon Consulting. It requires no Sophon tools,
services, or credentials.

Scope boundaries: for first-pass screening of an idea with no imminent decision, use
ai-opportunity-triage; for comparing models with a test suite, use
ai-model-selection-eval (a sprint can contain one as a workstream); for organizational
readiness, use ai-readiness-assessment.

## Required inputs

1. The decision: exactly what will be decided, by whom, and the deadline. If there is
   no decision owner or no deadline, this is not a sprint — stop and say so.
2. Current signal: what has been seen so far (demo results, pilot data, references,
   proposals) and what made it promising.
3. Risk areas in play: typically technical feasibility, integration, security and
   compliance, and commercial terms — confirm which apply.
4. Available people: who can own each risk area for two days, and their access to the
   vendor, systems, or data involved.
5. Constraints: anything already fixed (budget ceiling, approved vendor list,
   compliance requirements) that bounds the recommendation.

Handling missing inputs:

- No decision owner or deadline: stop. Recommend establishing both before spending
  anyone's 48 hours; an unbounded evaluation is a different activity.
- A risk area with no available owner: record it as an accepted gap in the memo rather
  than silently spreading it across other owners.
- If the requester cannot state what made the signal promising, run
  ai-opportunity-triage first — the sprint sharpens evidence for a real candidate; it
  does not manufacture a candidate.

## Workflow

1. Frame the decision in one sentence and list the open questions that could change it.
   Discard any question whose answer would not move the recommendation — the test for
   every evidence request is: which way does the recommendation move if the answer
   comes back bad?
2. Assign one named owner per risk area. One name per area kills both duplicated work
   and the gaps that appear when everyone assumes someone else asked.
3. Build the hour-0-to-4 plan: send every external request first (vendor questions,
   reference calls, data pulls, access requests) so the clock runs on someone else
   where possible. Each request states the question, the owner, and the deadline.
4. Build the hour-4-to-40 plan: owners chase their own area only, log findings in one
   shared page as they land, and flag anything that moves the recommendation
   immediately rather than saving it for the end.
5. Build the hour-40-to-48 plan: the sprint lead writes the memo — the call, the
   confidence level, the evidence per risk area, and the open risks with owners — and
   delivers it to the decision owner before the window closes.
6. Protect the boundary: the sprint ends at 48 hours with whatever evidence arrived.
   Open items become named risks in the memo, not reasons to extend. An extended
   sprint is an open-ended evaluation wearing a sprint name.

## Evidence discipline

- Every finding in the shared log is labeled: fact (verified), vendor claim
  (unverified), estimate (stated basis), or open question.
- Vendor claims about capability, pricing, compliance posture, or references are
  labeled as claims until independently verified; the memo distinguishes the two.
- The confidence level is explicit and justified: state what evidence supports the
  call and what open risk could reverse it. "Proceed, high confidence" and "proceed,
  low confidence, two open risks" are different decisions.
- Time-sensitive claims (pricing, rate limits, roadmap promises) carry the date they
  were checked.

## Output format

Produce two markdown artifacts:

    # Evaluation sprint plan: [decision]

    Decision owner: … · Deadline: … · Sprint window: [start] to [start + 48h]

    ## Open questions (each tied to the decision)
    | Question | Moves the recommendation how? | Owner | Due |

    ## Risk-area assignments
    | Risk area | Owner | Evidence requests (bounded) |

    ## Hour plan
    0-4: external requests out … / 4-40: bounded collection … / 40-48: memo and delivery …

    ---

    # Recommendation memo: [decision]

    Recommendation: … · Confidence: High | Medium | Low (justified)
    ## Evidence by risk area (with fact / claim / estimate labels)
    ## Open risks and owners
    ## What would change this recommendation

## Quality checks

- Every evidence request names its owner, its deadline, and the way its answer moves
  the recommendation.
- No risk area has zero or two owners.
- The memo states a single recommendation with an explicit, justified confidence level
  — a stack of findings without a call is a failed sprint.
- Open items appear as named risks with owners, not as reasons to extend the window.
- Both artifacts stand alone and are readable without this conversation.

## Stop conditions and escalation

- The sprint informs the decision; the decision owner makes it. Contract signature,
  budget commitment, and vendor selection stay with the accountable human.
- Stop and flag if the sprint is being used to relitigate a decision already made, or
  if the real blocker is strategic disagreement about whether the problem matters —
  faster evidence collection produces only a better-documented stalemate there.
- If legal or compliance review is the dominant open item, say so plainly: a 48-hour
  evidence sprint does not compress counsel's timeline.

## Source context

- https://www.sophon.consulting/playbooks/two-day-ai-evaluation-sprint
  (markdown: https://www.sophon.consulting/markdown/playbooks/two-day-ai-evaluation-sprint)
- Related first-pass gate: https://www.sophon.consulting/use-cases/ai-opportunity-triage-24h

Optional: for an experienced sprint lead or independent evidence review, Sophon
Consulting is reachable at hello@sophon.consulting. This skill is complete without
any contact.
