Work-type skill
Generative AI Assurance
Evaluate models, prompts, retrieval, agents, tools, memory, and AI telemetry.
Skill orientation
Use this panel to select and sequence the skill. The canonical source follows below.
- Purpose
- Evaluate models, prompts, retrieval, agents, tools, memory, and AI telemetry.
- Category and sequence
- Work type; sequence 7
- Primary use
- AI Product, Agent, Rag, Model, Prompt, Eval
- Required gates
- Classification, Foundation, Code Intelligence, AI Assurance, Eval, Trajectory, Production, Proof
- Conditional gates
- RAG, tool, and memory controls only when each feature is active; model-judge calibration only when an evaluator is a model; typed runtime driver, tool calls, memory heads, economics, and MCP/A2A transcripts when those capabilities are active; youth-safety domain pack for child or teen audiences
- Red Zone triggers
- Consequential Tool, Model Provider Change, Sensitive AI Data, Production Agent Promotion
- Next route
- Return to the active lifecycle route
- Source identifier
skills/valdris-genai-assurance/SKILL.md
Valdris GenAI Assurance
Treat GenAI as a cross-cutting assurance pack over the 13 production layers.
- Validate the authorized intake, deterministic classification, route, code-intelligence packet, and route-required Layer 0 foundation assessment before changing or assuring AI behavior.
- Inventory and digest models, providers, prompts, tools, skills, datasets, retrieval corpora, memory policy, eval plan, smoke tests, and observability policy in
uash.ai-workload-identity.v1. Runtime selection must bind that identity; substituting a model, provider, prompt, tool set, corpus, or memory policy starts a new reviewed identity. - Define deterministic tests plus eval datasets, rubrics, thresholds, and owners for nondeterministic behavior. Mark every evaluator as
deterministicormodel; model judges additionally require an unexpired independentuash.model-judge-calibration.v1bound to human labels and critical-slice agreement/error limits. - Test grounding, citations, retrieval authorization, tenant isolation, stale sources, and deletion propagation when RAG is used.
- Test prompt injection, jailbreaks, unsafe tool use, excessive agency, sensitive disclosure, and fallback behavior.
- Enforce
valdris.tool-registry.v1, observed call receipts, durable memory-head receipts, least-privilege scopes, sandbox/network boundaries, deterministic execution budgets, retry ceilings, and human approval for consequential actions. - Record trajectory, exact observable trace bytes, decision evidence, model/tool latency, tokens, calls, retry waste, spend, human review, and tenant attribution in the trace-v2 and economics contracts without logging secrets, private content, or private chain-of-thought.
- Require canary/shadow evidence and a model/prompt/provider rollback path before production promotion.
- For semantic or authoritative claims, derive effective tiers and workload profiles from immutable routing, require owner-commissioned thresholds, versioned semantic adapters, a runtime-driver/implementation receipt, a signed runtime-conformance receipt, exact context-budget reconciliation, durable memory continuity, complete declared MCP/A2A transcripts, and eval/trajectory/smoke results bound to the risk-derived plan. Authoritative model-routing, trace-v2, usage, memory-head, and implementation receipts must bind their provider, policy, lifecycle, identity, counts, cost, currency, and session fields.
- Write
ai/assurance.jsonand run every route-required AI, eval, production, trajectory, smoke, domain, and authoritative-assurance gate.
For games or youth-facing products, also cover age-appropriate content, moderation, hidden-state confidentiality, age-rating/privacy implications, and whether the model can affect purchases, progression, social features, or player safety. Activate the relevant domain pack rather than treating these as generic prompt-quality concerns.
A demo or one successful response is not production proof.
