Enterprise deep dives¶
The six labs, including model routing and failover, demonstrate individual controls. A real deployment also needs durable identity, policy lifecycle, data protection, reliability, and evidence management.
Production reference path¶
| Concern | Demo behavior | Enterprise enhancement |
|---|---|---|
| Agent identity | Local/example JWT | Workload identity, short-lived tokens, issuer/audience validation, key rotation |
| Authorization | Static CEL rules | Policy-as-code repository, owners, review gates, versioned decisions, deny-by-default |
| Tool risk | Input patterns and description guardrails | Tool registry, risk classification, output filtering, egress controls, threat modeling |
| Audit | Logs and traces | Immutable evidence store, retention policy, redaction, identity-to-trace correlation |
| Workflow state | In-memory task map | Durable state, idempotency keys, optimistic concurrency, timeout and compensation |
| Approval | Demo admin endpoint | Authenticated approver identity, separation of duties, reason, expiry, replay prevention |
| Evaluation | Deterministic golden files | Versioned datasets, environment matrix, policy negative tests, trend and release thresholds |
| Model routing | Local stubs and bounded fallback | Approved providers/regions, credential management, streaming and retry tests, model-quality evaluation |
Read the production-readiness guide for a concrete architecture, trust boundaries, failure modes, and a staged backlog.
Recommended technical deep-dive labs¶
These are the highest-value next examples for this repository:
- OIDC workload identity and key rotation: replace the sample JWT with a real issuer, JWKS caching, audience checks, expiry tests, and a rotation drill.
- Tenant isolation: carry
tenant_idfrom token to policy, traces, cache keys, and backend queries; prove cross-tenant requests fail. - Durable HITL workflow: persist tasks and approvals, add idempotency, expiry, concurrency tests, and an append-only decision journal.
- Telemetry privacy: redact PII at source, enforce attribute allowlists, sample by risk, and test that secrets never reach the collector.
- Adversarial evaluation: test confused-deputy calls, prompt/tool poisoning, replay, oversized payloads, authorization drift, and approval bypass.