Skip to harness content
Open technical reference map

Work-type skill

Generative AI Assurance

Evaluate models, prompts, retrieval, agents, tools, memory, and AI telemetry.

valdris-genai-assuranceOwning system: Active lifecycle owner

Skill orientation

Use this panel to select and sequence the skill. The canonical source follows below.

Purpose
Evaluate models, prompts, retrieval, agents, tools, memory, and AI telemetry.
Category and sequence
Work type; sequence 7
Primary use
AI Product, Agent, Rag, Model, Prompt, Eval
Required gates
Classification, Foundation, Code Intelligence, AI Assurance, Eval, Trajectory, Production, Proof
Conditional gates
RAG, tool, and memory controls only when each feature is active; model-judge calibration only when an evaluator is a model; typed runtime driver, tool calls, memory heads, economics, and MCP/A2A transcripts when those capabilities are active; youth-safety domain pack for child or teen audiences
Red Zone triggers
Consequential Tool, Model Provider Change, Sensitive AI Data, Production Agent Promotion
Next route
Return to the active lifecycle route
Source identifier
skills/valdris-genai-assurance/SKILL.md
Canonical pathskills/valdris-genai-assurance/SKILL.mdRevision69bab1cInspect source

Valdris GenAI Assurance

Treat GenAI as a cross-cutting assurance pack over the 13 production layers.

  1. Validate the authorized intake, deterministic classification, route, code-intelligence packet, and route-required Layer 0 foundation assessment before changing or assuring AI behavior.
  2. Inventory and digest models, providers, prompts, tools, skills, datasets, retrieval corpora, memory policy, eval plan, smoke tests, and observability policy in uash.ai-workload-identity.v1. Runtime selection must bind that identity; substituting a model, provider, prompt, tool set, corpus, or memory policy starts a new reviewed identity.
  3. Define deterministic tests plus eval datasets, rubrics, thresholds, and owners for nondeterministic behavior. Mark every evaluator as deterministic or model; model judges additionally require an unexpired independent uash.model-judge-calibration.v1 bound to human labels and critical-slice agreement/error limits.
  4. Test grounding, citations, retrieval authorization, tenant isolation, stale sources, and deletion propagation when RAG is used.
  5. Test prompt injection, jailbreaks, unsafe tool use, excessive agency, sensitive disclosure, and fallback behavior.
  6. Enforce valdris.tool-registry.v1, observed call receipts, durable memory-head receipts, least-privilege scopes, sandbox/network boundaries, deterministic execution budgets, retry ceilings, and human approval for consequential actions.
  7. Record trajectory, exact observable trace bytes, decision evidence, model/tool latency, tokens, calls, retry waste, spend, human review, and tenant attribution in the trace-v2 and economics contracts without logging secrets, private content, or private chain-of-thought.
  8. Require canary/shadow evidence and a model/prompt/provider rollback path before production promotion.
  9. For semantic or authoritative claims, derive effective tiers and workload profiles from immutable routing, require owner-commissioned thresholds, versioned semantic adapters, a runtime-driver/implementation receipt, a signed runtime-conformance receipt, exact context-budget reconciliation, durable memory continuity, complete declared MCP/A2A transcripts, and eval/trajectory/smoke results bound to the risk-derived plan. Authoritative model-routing, trace-v2, usage, memory-head, and implementation receipts must bind their provider, policy, lifecycle, identity, counts, cost, currency, and session fields.
  10. Write ai/assurance.json and run every route-required AI, eval, production, trajectory, smoke, domain, and authoritative-assurance gate.

For games or youth-facing products, also cover age-appropriate content, moderation, hidden-state confidentiality, age-rating/privacy implications, and whether the model can affect purchases, progression, social features, or player safety. Activate the relevant domain pack rather than treating these as generic prompt-quality concerns.

A demo or one successful response is not production proof.