# Engineering Runbook (PEDW FrontEnd) ## Purpose Provide a repeatable operational playbook for safe delivery and incident-aware change management. > Pre-flight: run `GUARDRAILS.md` checklist before significant implementation or merge. ## Standard Change Workflow 1. Clarify scope, constraints, and affected flows (public search/case, account/auth, portal/admin). 2. Identify risk level: - **High:** auth/session, Prisma schema, middleware/CSP, uploads/documents, notifications - **Medium:** routing/i18n rewrites, Redux hydration/persistence, PDF generation - **Low:** isolated UI or content-only updates 3. Implement smallest viable change with clear rollback path. 4. Validate (minimum): - `npm run lint` - targeted manual checks for touched routes/APIs - EN/CY parity checks - accessibility smoke (keyboard, focus, labels, headings) - negative-path checks for sensitive logic 5. Record outcomes in `memory-bank/change-log.md`. ## Pre-Release Checklist - [ ] Scope and non-goals documented - [ ] Risk notes cover auth/data/i18n/a11y impacts - [ ] Validation evidence captured (command output + manual matrix) - [ ] Rollback steps prepared - [ ] Memory-bank updated for decisions/pitfalls/open questions ## Incident Triage (Quick) 1. Classify impact: service unavailable, degraded flow, data/security concern. 2. Identify blast radius: which routes/APIs/locales/users are affected. 3. Check telemetry and logs (server + app insights) for errors around deploy window. 4. Apply mitigation: - rollback/revert high-risk change, - or ship minimal hotfix with guarded scope. 5. Validate recovery on EN/CY critical journeys. 6. Capture post-incident notes in `memory-bank/change-log.md` and `memory-bank/pitfalls.md`. ## High-Risk Flow Verification Matrix - **Auth:** sign-in, verify-request, callback redirect safety, session continuity - **Portal/account:** protected route access and profile/account updates - **Uploads/docs:** invalid file type/size handling, unauthorized access protection - **Notifications:** template/language selection and failure handling - **Routing/i18n:** rewritten Welsh routes land on expected handlers ## Relay Hardening Rollout Playbook (TASK22239) Use this when changing shared relay forwarding policy (timeouts/retries/logging) or deploying relay policy updates. ### Pre-Merge Governance Gate - [ ] PR scope states what changed and what did not (endpoint contracts vs relay internals) - [ ] Risk notes include auth/data/i18n/a11y impact and relay behavior impact - [ ] Validation evidence attached: - relay hardening targeted tests - endpoint contract regression tests - lint status - [ ] Rollback steps documented (config rollback + commit revert) - [ ] Memory-bank updated (`change-log`, and where relevant `decisions`/`patterns`) ### Non-Prod Smoke Matrix (Required) Run against a representative non-production environment: 1. **Deterministic auth/client failures** - force/verify `401` and `403` - expected: no retries, immediate handled failure 2. **Deterministic validation failures** - force/verify `400` or `404` - expected: no retries 3. **Transient upstream failures** - force/verify `503` / `429` - expected: bounded retries + bounded backoff 4. **Timeout behavior** - force latency above timeout threshold - expected: bounded failure path and no retry storm 5. **Operational logging behavior** - expected: structured redacted retry/failure events - expected: no duplicate endpoint-layer error spam for already-logged relay failures ### Progressive Runtime Rollout 1. Deploy with conservative retry settings. 2. Verify service stability and log volume for first release window. 3. Tune only one variable at a time (`timeout`, then `retry count`, then delays). Recommended starting posture: - `RELAY_RETRY_MAX` in low range (e.g. `1` or `2`) - bounded delay values aligned with user-facing latency tolerance - avoid simultaneous increases of timeout and retries unless incident evidence requires it ### Monitoring Checks (Day 1 / Day 3) - Relay retry rate trend - Timeout/error ratio and top status buckets - Upstream latency impact on citizen-facing journeys - Log volume increase/decrease and duplicate-error noise ### Fast Rollback / Mitigation 1. Immediate mitigation: set `RELAY_RETRY_MAX=0` (disables retries without code rollback). 2. If needed, reduce timeout and delay knobs to baseline-safe values. 3. Full rollback path: revert relay hardening commit set and redeploy. 4. Record incident + mitigation outcome in `memory-bank/change-log.md` and `memory-bank/pitfalls.md`. ## Definition of Ready for AI-Assisted Tasks - Clear acceptance criteria - Named affected files/flows - Risk classification assigned - Validation plan agreed (including manual checks)