119 lines
4.7 KiB
Markdown
119 lines
4.7 KiB
Markdown
# Engineering Runbook (PEDW FrontEnd)
|
|
|
|
## Purpose
|
|
|
|
Provide a repeatable operational playbook for safe delivery and incident-aware change management.
|
|
|
|
> Pre-flight: run `GUARDRAILS.md` checklist before significant implementation or merge.
|
|
|
|
## Standard Change Workflow
|
|
|
|
1. Clarify scope, constraints, and affected flows (public search/case, account/auth, portal/admin).
|
|
2. Identify risk level:
|
|
- **High:** auth/session, Prisma schema, middleware/CSP, uploads/documents, notifications
|
|
- **Medium:** routing/i18n rewrites, Redux hydration/persistence, PDF generation
|
|
- **Low:** isolated UI or content-only updates
|
|
3. Implement smallest viable change with clear rollback path.
|
|
4. Validate (minimum):
|
|
- `npm run lint`
|
|
- targeted manual checks for touched routes/APIs
|
|
- EN/CY parity checks
|
|
- accessibility smoke (keyboard, focus, labels, headings)
|
|
- negative-path checks for sensitive logic
|
|
5. Record outcomes in `memory-bank/change-log.md`.
|
|
|
|
## Pre-Release Checklist
|
|
|
|
- [ ] Scope and non-goals documented
|
|
- [ ] Risk notes cover auth/data/i18n/a11y impacts
|
|
- [ ] Validation evidence captured (command output + manual matrix)
|
|
- [ ] Rollback steps prepared
|
|
- [ ] Memory-bank updated for decisions/pitfalls/open questions
|
|
|
|
## Incident Triage (Quick)
|
|
|
|
1. Classify impact: service unavailable, degraded flow, data/security concern.
|
|
2. Identify blast radius: which routes/APIs/locales/users are affected.
|
|
3. Check telemetry and logs (server + app insights) for errors around deploy window.
|
|
4. Apply mitigation:
|
|
- rollback/revert high-risk change,
|
|
- or ship minimal hotfix with guarded scope.
|
|
5. Validate recovery on EN/CY critical journeys.
|
|
6. Capture post-incident notes in `memory-bank/change-log.md` and `memory-bank/pitfalls.md`.
|
|
|
|
## High-Risk Flow Verification Matrix
|
|
|
|
- **Auth:** sign-in, verify-request, callback redirect safety, session continuity
|
|
- **Portal/account:** protected route access and profile/account updates
|
|
- **Uploads/docs:** invalid file type/size handling, unauthorized access protection
|
|
- **Notifications:** template/language selection and failure handling
|
|
- **Routing/i18n:** rewritten Welsh routes land on expected handlers
|
|
|
|
## Relay Hardening Rollout Playbook (TASK22239)
|
|
|
|
Use this when changing shared relay forwarding policy (timeouts/retries/logging) or deploying relay policy updates.
|
|
|
|
### Pre-Merge Governance Gate
|
|
|
|
- [ ] PR scope states what changed and what did not (endpoint contracts vs relay internals)
|
|
- [ ] Risk notes include auth/data/i18n/a11y impact and relay behavior impact
|
|
- [ ] Validation evidence attached:
|
|
- relay hardening targeted tests
|
|
- endpoint contract regression tests
|
|
- lint status
|
|
- [ ] Rollback steps documented (config rollback + commit revert)
|
|
- [ ] Memory-bank updated (`change-log`, and where relevant `decisions`/`patterns`)
|
|
|
|
### Non-Prod Smoke Matrix (Required)
|
|
|
|
Run against a representative non-production environment:
|
|
|
|
1. **Deterministic auth/client failures**
|
|
- force/verify `401` and `403`
|
|
- expected: no retries, immediate handled failure
|
|
2. **Deterministic validation failures**
|
|
- force/verify `400` or `404`
|
|
- expected: no retries
|
|
3. **Transient upstream failures**
|
|
- force/verify `503` / `429`
|
|
- expected: bounded retries + bounded backoff
|
|
4. **Timeout behavior**
|
|
- force latency above timeout threshold
|
|
- expected: bounded failure path and no retry storm
|
|
5. **Operational logging behavior**
|
|
- expected: structured redacted retry/failure events
|
|
- expected: no duplicate endpoint-layer error spam for already-logged relay failures
|
|
|
|
### Progressive Runtime Rollout
|
|
|
|
1. Deploy with conservative retry settings.
|
|
2. Verify service stability and log volume for first release window.
|
|
3. Tune only one variable at a time (`timeout`, then `retry count`, then delays).
|
|
|
|
Recommended starting posture:
|
|
|
|
- `RELAY_RETRY_MAX` in low range (e.g. `1` or `2`)
|
|
- bounded delay values aligned with user-facing latency tolerance
|
|
- avoid simultaneous increases of timeout and retries unless incident evidence requires it
|
|
|
|
### Monitoring Checks (Day 1 / Day 3)
|
|
|
|
- Relay retry rate trend
|
|
- Timeout/error ratio and top status buckets
|
|
- Upstream latency impact on citizen-facing journeys
|
|
- Log volume increase/decrease and duplicate-error noise
|
|
|
|
### Fast Rollback / Mitigation
|
|
|
|
1. Immediate mitigation: set `RELAY_RETRY_MAX=0` (disables retries without code rollback).
|
|
2. If needed, reduce timeout and delay knobs to baseline-safe values.
|
|
3. Full rollback path: revert relay hardening commit set and redeploy.
|
|
4. Record incident + mitigation outcome in `memory-bank/change-log.md` and `memory-bank/pitfalls.md`.
|
|
|
|
## Definition of Ready for AI-Assisted Tasks
|
|
|
|
- Clear acceptance criteria
|
|
- Named affected files/flows
|
|
- Risk classification assigned
|
|
- Validation plan agreed (including manual checks)
|