Pouya's four rulings of 2026-09-01, applied 2026-09-02.
Q62 — RULED "state the truth", not "remove the access". /legal/privacy/
now says two people can read the intake table, names their role, and adds
the two stronger facts the false sentence had crowded out: the handler
role holds PutItem only, and adr-sml-deploy is implicitDeny on all seven
read and write actions. Wording is subject to Pouya's read-through —
Q63(a), with a TODO(pouya) beside the copy.
The ruling named one sentence; a vocabulary sweep found the falsehood in
three places, and the audit then found two more. Five paragraphs now
answer "who can see it" and change together.
The tripwire stays permanently, per ruling, and grew from two
alternatives to five. Every alternative is one string that reached dist/.
Proven both ways against the pre-correction page rebuilt from bd282aa:
exit 1 with 5 matches at dist/legal/privacy/index.html:54,67,67,68,72;
exit 0 on the corrected page, self-test 12 patterns / 36 approved
strings.
/med-arb/ — the gloss is struck with no replacement, per ruling. The
strike left "the section above" pointing at the ADRIC rule set and "the
agreement" with no antecedent; both fixed. The bare designations line
sitting under ADRIC's quoted competence requirement is also struck, which
goes beyond the ruling and is flagged for Pouya.
R18 — re-stamped, two-tier: (a)(c)(d) re-verified against a source,
(b)(e)(f)(g) held on a cadence judgement. All seven hold, no shipped
sentence changed. R18's trigger had NO cutover checklist item and had
stamped five extracts of seven; both fixed. Candidate limb (h) flagged.
R10 — fired and unsatisfied; left open on instruction.
The evidence behind the new privacy sentence was weaker than the
sentence. Re-measured: 33 of 33 roles simulated (23 of 26 carried inline
policies nobody had read; the two CDK lookup roles can read the table),
four trust policies, the CloudFormation escalation path for all five
users, 0 federated providers, root recorded. Every read path terminates
at the same two people.
Two review rounds, 36 findings. 35 fixed, 1 declined. Five of round 2's
were defects in round 1's own fixes; stopped at two per D19.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Md3GndFqWPzK78xAoebsg5
47 KiB
09 — Cutover runbook: the exact command sequence
Authority: AGENTS.md §3 D11 (build everything, one clean cutover) and §7 for
every operational fact. docs/06-deployment.md owns the cutover checklist —
what must be true. This file owns the commands — how to make it true. The
checklist cites this file; this file does not restate the checklist.
Pouya runs everything here. Parts 1–7 need administrative credentials that no
agent on this project holds, and Part 8 is a deploy that must not run as
user/pouya (§10). Every command is followed by the command that verifies it and
the output to expect. Where a verification can come back two ways, both are named
and the sequence branches — a branch driven by a measurement rather than by a
guess is the point.
Read
docs/06's cutover checklist alongside this. The parts below discharge its Technical group. Its Content and compliance group — theclaims-auditorpass and Pouya's own page-by-page read — is not commands and is not here.
Why the order is this order, and what is broken while it runs
The bucket currently holds the old single-file site: index.html (2,206,032
bytes) plus two logo PNGs [verified 2026-09-01 — list-objects-v2]. So the live
site today is one page, and one page is all that works: measured on the live
distribution the same day, / returns 200 while /about/ and
/definitely-not-a-page/ both return 403 with an 111-byte application/xml
body — S3's AccessDenied, served raw to the reader.
That is why the infrastructure goes first and the site goes last:
- Parts 1–4 (distribution) change nothing a visitor can see. The router
rewrites
/to/index.html, which is what the default root object already did, and the paths it newly handles do not exist on the old site either way. - Parts 5–7 (intake) are inert until
/contact/exists. - Part 8 is the cutover. The moment
index.htmlsyncs, the new site is live — and by then the router, the 404 mapping and/api/*are already in place, so there is no window in which 22 of 23 pages are broken.
Reversing this — site first, infrastructure second — would put every page except
/ behind a 403 for as long as the CloudFront deployment took.
Part 0 — Preflight
0.1 The substitution block
⚠️ THIS IS THE ONLY PLACE IN THIS FILE THAT NAMES A RESOURCE, AND THAT IS A
DELIBERATE, BOUNDED EXCEPTION TO THE §7 RULE. AGENTS.md §7 is the single
source of truth for operational facts and specs cite it rather than copying it —
but a runbook that cannot be pasted into a terminal is not a runbook. So the
values appear exactly once, and 0.2 is the mechanism that keeps the copy
honest: it proves every one of them resolves to a real resource before anything
is changed. A duplicated fact needs a mechanism, not a promise.
Confirm each against §7 before pasting. ./scripts/aws-discover.sh regenerates
them from AWS if §7 itself needs re-verifying.
cd ~/Dev/Websites/adr-sml
export AWS_REGION=ca-central-1
export AWS_DEFAULT_REGION=ca-central-1
export ACCT=327082975128
export FN=adr-intake-handler
export TABLE=adr-intake-submissions
export API_ID=4tl0m5igkj
export API_DOMAIN=4tl0m5igkj.execute-api.ca-central-1.amazonaws.com
export DIST_ID=E1OK7G98KNKUTA
export BUCKET=adr-smlcompany-site
export SITE=https://adr.smlcompany.ca
export NOTIFY_TO=info@smlcompany.ca
export MAIL_FROM=intake@smlcompany.ca
NOTIFY_TO is where an inquiry lands and MAIL_FROM is the From on both
messages. Both are choices rather than records: info@ is the address /contact/
publishes, so a reply threads where the inquirer expects; intake@ distinguishes
form mail from correspondence. Both must be at smlcompany.ca, which is the
verified sending identity — 0.2 checks that.
0.2 Prove every one of them exists
aws sts get-caller-identity
for v in AWS_REGION ACCT FN TABLE API_ID API_DOMAIN DIST_ID BUCKET SITE NOTIFY_TO MAIL_FROM; do
eval "val=\$$v"
[ -n "$val" ] || { echo "EMPTY: $v"; break; }
printf '%-12s %s\n' "$v" "$val"
done
aws lambda get-function-configuration --function-name "$FN" --query 'FunctionName' --output text
aws dynamodb describe-table --table-name "$TABLE" --query 'Table.TableStatus' --output text
aws apigatewayv2 get-api --api-id "$API_ID" --query 'Name' --output text
aws cloudfront get-distribution --id "$DIST_ID" --query 'Distribution.Status' --output text
aws s3api head-bucket --bucket "$BUCKET" && echo "bucket ok"
aws sesv2 get-email-identity --email-identity "${MAIL_FROM#*@}" \
--query 'VerifiedForSendingStatus' --output text
Expect: an identity with administrative rights; eleven non-empty values;
adr-intake-handler; ACTIVE; adr-intake-api; Deployed; bucket ok; True.
A None, an empty value, or any non-zero exit stops the run here. That is the
whole purpose of 0.2 — a stale identifier fails loudly at the top instead of
halfway through Part 6.
0.3 State this run starts from
Recorded so a later reader can tell what this sequence changed from what it found.
All [verified 2026-09-01], read-only:
| Lambda | nodejs24.x, arm64, handler index.handler, timeout 10 s, memory 128 MB, no environment variables, no DLQ, code 1,527 bytes (the hand-built inline function), last modified 2026-05-26 |
| Lambda role | adr-intake-lambda-role — AWSLambdaBasicExecutionRole plus an inline policy granting dynamodb:PutItem on the table and ses:SendEmail/ses:SendRawEmail. Sufficient for the new handler; no change needed |
| Lambda invoke permission | one statement, apigateway.amazonaws.com, SourceArn scoped to …/4tl0m5igkj/*/*/submissions — the old route's path only |
| HTTP API | one route, POST /submissions → integration 0ftgjgv (AWS_PROXY, payload format 2.0, the correct one for this handler). Stage $default, auto-deploy on, no throttling, no access log. CORS allows POST from the site origin |
| DynamoDB | ACTIVE, on-demand, partition key submissionId (S), no sort key, 4 existing items, TTL ENABLED on attribute ttl, PITR ENABLED (35-day window), encryption at rest with the AWS-owned key (no CMK) |
| CloudFront | one S3 origin with OAC E13GAFUL6UQP6R; default behaviour Managed-CachingOptimized + Managed-SecurityHeadersPolicy, methods HEAD/GET; no cache behaviours, no custom error responses, no function associations |
| S3 | versioning Enabled; all four public-access blocks true; direct object GET on the S3 hostname returns 403, so OAC is in force |
| SES | domain verified for sending, DKIM SUCCESS and signing enabled, no custom MAIL FROM |
SNS ses-alerts |
email subscription to info@smlcompany.ca is CONFIRMED — it has a real subscription ARN, not PendingConfirmation. §7 recorded it as pending for six days after it stopped being pending; §12 R9 closes on this |
Part 1 — Make S3 answer 404 instead of 403
Why this is first. Part 3 maps 404 to /404.html and deliberately does
not map 403. With OAC and no s3:ListBucket, S3 answers a request for a
missing key with 403 AccessDenied, so a 404-only mapping would never fire and
every bad URL would keep returning raw XML. Granting the CloudFront service
principal s3:ListBucket changes that answer to 404 NoSuchKey.
Mapping 403 as well would have been one command shorter and is the wrong trade:
custom error responses are distribution-wide and cannot be scoped to one
behaviour, so it would also convert a broken bucket policy or a broken OAC — a
real outage on every URL at once — into a tidy "page not found", and it would
turn the intake handler's Origin refusal into a 404 page. Keeping 403 meaning
403 is worth one statement.
aws s3api get-bucket-policy --bucket "$BUCKET" --query Policy --output text > /tmp/bucket-policy.before.json
python3 -m json.tool /tmp/bucket-policy.before.json
Expect: one statement, AllowCloudFrontServicePrincipal, s3:GetObject on
arn:aws:s3:::$BUCKET/*, conditioned on the distribution ARN. Keep that file —
Part 9 restores from it.
python3 - "$BUCKET" "$ACCT" "$DIST_ID" <<'PY' > /tmp/bucket-policy.after.json
import json, sys
bucket, acct, dist = sys.argv[1], sys.argv[2], sys.argv[3]
p = json.load(open('/tmp/bucket-policy.before.json'))
arn = f'arn:aws:cloudfront::{acct}:distribution/{dist}'
sid = 'AllowCloudFrontListForHonest404s'
p['Statement'] = [s for s in p['Statement'] if s.get('Sid') != sid]
p['Statement'].append({
'Sid': sid,
'Effect': 'Allow',
'Principal': {'Service': 'cloudfront.amazonaws.com'},
'Action': 's3:ListBucket',
'Resource': f'arn:aws:s3:::{bucket}',
'Condition': {'StringEquals': {'AWS:SourceArn': arn}},
})
json.dump(p, sys.stdout, indent=2)
PY
python3 -m json.tool /tmp/bucket-policy.after.json
aws s3api put-bucket-policy --bucket "$BUCKET" --policy "file:///tmp/bucket-policy.after.json"
Verify — three things, and the third is the one worth doing:
aws s3api get-bucket-policy --bucket "$BUCKET" --query Policy --output text | python3 -m json.tool
aws s3api get-public-access-block --bucket "$BUCKET" --query 'PublicAccessBlockConfiguration'
curl -s -o /dev/null -w 'direct S3 object: %{http_code}\n' \
"https://${BUCKET}.s3.${AWS_REGION}.amazonaws.com/index.html"
aws cloudfront get-distribution-config --id "$DIST_ID" \
--query 'DistributionConfig.DefaultCacheBehavior.{Cache:CachePolicyId,OriginRequest:OriginRequestPolicyId}'
aws cloudfront get-cache-policy --id 658327ea-f89d-4fab-a63d-7e88639e58f6 \
--query 'CachePolicy.CachePolicyConfig.ParametersInCacheKeyAndForwardedToOrigin.QueryStringsConfig'
Expect: two statements; all four public-access blocks still true; direct S3
403; Cache = 658327ea-… (Managed-CachingOptimized) with
OriginRequest: null; and QueryStringsConfig = {"QueryStringBehavior": "none"}".
⚠️ THOSE LAST TWO ARE THE CHECK, AND THE OBVIOUS ONE IS WORTHLESS. The
tempting verification is curl "${SITE}/?list-type=2&max-keys=5" and confirming
it returns HTML rather than an XML ListBucketResult — it cannot return
anything else, so it is evidence of nothing. The distribution has
DefaultRootObject: index.html and, after Part 2, the router rewrites / to
/index.html; the origin request is therefore a GetObject on a key, never a
request for the bucket root, so list-type=2 could not be read as a list
operation whether or not query strings reached S3. An operator running it sees
HTML and ticks a control that never ran, which is the Q22 shape exactly. Found by
adversarial-reviewer, 2026-09-01.
What actually makes the grant safe, and the two commands above assert both halves: no viewer path resolves to the bucket root, and the default behaviour forwards no query strings to the origin — with no origin request policy attached to override that. If either assertion fails, revert this part (Part 9.1) and map 403 instead.
Part 2 — The trailing-slash router function
Source and test are in the repo: infra/cloudfront/router.js and
router.test.mjs. Read the header comment before publishing it — the reason it
must not be associated with /api/* is in there, and it is the difference
between a working form and a POST silently converted to a GET.
node infra/cloudfront/router.test.mjs
Expect: exit 0. The case count is deliberately not quoted here — it has
already gone 15 → 22 → 30 in one day, and an Expect line that never matches
teaches the operator to stop reading Expect lines. The script asserts its own
count internally (case count != CASES.length), so exit 0 is the whole check.
aws cloudfront create-function \
--name adr-sml-router \
--function-config '{"Comment":"trailing-slash + index.html for the Astro directory build; docs/09","Runtime":"cloudfront-js-2.0"}' \
--function-code "fileb://infra/cloudfront/router.js" \
--query '{Name:FunctionSummary.FunctionMetadata.FunctionARN,Stage:FunctionSummary.FunctionMetadata.Stage,Status:FunctionSummary.Status}'
Expect: an ARN, Stage: DEVELOPMENT, Status: UNPUBLISHED.
Test it in the real runtime before publishing. The local test proves the branches; this proves the runtime accepts the code, which the local test cannot.
ETAG=$(aws cloudfront describe-function --name adr-sml-router --query ETag --output text)
for URI in / /about/ /about /robots.txt /_astro/x.css /api/intake \
//evil.example.com/x '/\evil.example.com/x' //about/ //robots.txt; do
printf '{"version":"1.0","request":{"method":"GET","uri":"%s","querystring":{},"headers":{},"cookies":{}}}' "$URI" \
> /tmp/cf-event.json
printf '%-16s ' "$URI"
aws cloudfront test-function --name adr-sml-router --if-match "$ETAG" --stage DEVELOPMENT \
--event-object fileb:///tmp/cf-event.json \
--query 'TestResult.{Out:FunctionOutput,Err:FunctionErrorMessage}' --output text
done
fileb:// rather than a base64 argument: it hands the CLI raw bytes and sidesteps
the question of whether the local base64 wraps its output, which would make the
argument invalid in a way the error message would not explain. test-function
does not change the function, so one ETAG read serves the whole loop.
Then the query-string path, which none of the URIs above exercises:
cat > /tmp/cf-event-qs.json <<'JSON'
{"version":"1.0","request":{"method":"GET","uri":"/fees","headers":{},"cookies":{},
"querystring":{"utm":{"value":"a%20b"},"q":{"value":"x|y"},"bad":{"value":"z\"<>#&k=v"}}}}
JSON
aws cloudfront test-function --name adr-sml-router --if-match "$ETAG" --stage DEVELOPMENT \
--event-object fileb:///tmp/cf-event-qs.json \
--query 'TestResult.{Out:FunctionOutput,Err:FunctionErrorMessage}' --output text
Expect, in order for the ten URIs: uri /index.html; uri /about/index.html; a 301 to /about/; uri /robots.txt unchanged; uri /_astro/x.css unchanged; /api/intake a 301 to /api/intake/; and then
four 301s that all stay on this origin with a single leading slash —
/evil.example.com/x/, /evil.example.com/x/, /about/, /robots.txt.
No FunctionErrorMessage on any of the ten.
From the query-string call: a 301 to
/fees/?utm=a%20b&q=x|y&bad=z&k=v. Three things are being checked there and
each has been wrong once: %20 survives (the values arrive percent-encoded, so
re-encoding would produce %2520); | survives (it is not in WHATWG's query
percent-encode set, so browsers send it raw, and one revision of the function
stripped it and silently corrupted campaign links); and ", <, > and # are
gone (# is the one that changes the structure of the header — left in, &k=v
lands in a fragment and the parameter is lost).
⚠️ THE FIRST VERSION OF THIS LOOP PASSED "querystring":{} ON EVERY CASE AND
TRIED NEITHER REDIRECT PATH. So in the only environment that runs the real
cloudfront-js-2.0 runtime — the environment this step exists to reach — none of
the query-string handling and none of the normalisation was executed, while the
prose called it "the authoritative check". Found by adversarial-reviewer
round 2.
⚠️ That last row is the reason this function must not be associated with the
/api/* behaviour, and it is in the list so you see it rather than read about
it. A 301 turns a POST into a GET, so an intake submission routed through this
function would arrive at the handler as a GET with no body — and the visitor would
be told nothing was wrong. Part 3 associates the function with the default
behaviour only; configure.mjs sets FunctionAssociations: { Quantity: 0 } on
the /api/* behaviour explicitly rather than by omission.
aws cloudfront publish-function --name adr-sml-router --if-match "$ETAG"
export ROUTER_ARN=$(aws cloudfront describe-function --name adr-sml-router --stage LIVE \
--query 'FunctionSummary.FunctionMetadata.FunctionARN' --output text)
echo "ROUTER_ARN=$ROUTER_ARN"
Expect: a non-empty ARN. It is empty if publish-function failed — read the
status, not the absence of an error.
Part 3 — Apply the three distribution changes
One script, infra/cloudfront/configure.mjs, because the alternative is
hand-editing a 300-line JSON document and posting it back with an IfMatch ETag.
It is dry-run by default, idempotent, and it resolves both managed policy
ids by name from the account rather than carrying them as literals.
node infra/cloudfront/configure.mjs --dist "$DIST_ID" --api-domain "$API_DOMAIN" \
--function-arn "$ROUTER_ARN"
Expect — this is the dry run, and the output on a distribution in the state Part 0.3 records is exactly:
resolved Managed-CachingDisabled = 4135ea2d-6df8-44a3-9df3-4b5a84be39ad
resolved Managed-AllViewerExceptHostHeader = b689b0a8-53d0-40ab-baf2-68738e2966ac
4 change(s) to distribution E1OK7G98KNKUTA (ETag …):
+ DefaultCacheBehavior.FunctionAssociations viewer-request -> arn:…:function/adr-sml-router
+ CustomErrorResponses += 404 -> /404.html with status 404
+ Origins += intake-api -> …execute-api… (https-only, TLSv1.2)
+ CacheBehaviors += /api/* -> intake-api, CachingDisabled, AllViewerExceptHostHeader, POST allowed
DRY RUN — nothing was sent. Re-run with --apply to write it.
Fewer than four changes means part of this is already done — read which lines are
prefixed · (already present) and carry on. More than four, or a different set,
means the distribution is not in the state 0.3 recorded: stop and re-read it.
node infra/cloudfront/configure.mjs --dist "$DIST_ID" --api-domain "$API_DOMAIN" \
--function-arn "$ROUTER_ARN" --apply
aws cloudfront wait distribution-deployed --id "$DIST_ID"
echo "deployed: $?"
Verify the config landed:
aws cloudfront get-distribution-config --id "$DIST_ID" \
--query 'DistributionConfig.{Fn:DefaultCacheBehavior.FunctionAssociations.Items[].EventType,Err:CustomErrorResponses.Items[].{Code:ErrorCode,Page:ResponsePagePath,Status:ResponseCode},Beh:CacheBehaviors.Items[].{P:PathPattern,O:TargetOriginId,Methods:AllowedMethods.Items},Origins:Origins.Items[].Id}'
Expect: Fn: ["viewer-request"]; one error response 404 → /404.html → 404;
one behaviour /api/* → intake-api with POST in its method list; two origins.
Part 4 — Verify the distribution before the site exists
Run this now, against the old bucket contents. It is the check that the plumbing works independently of the deploy.
for p in / /index.html /about/ /nope-not-a-page/; do
printf '%-22s ' "$p"
curl -s -o /dev/null -w 'status=%{http_code} ct=%{content_type}\n' "${SITE}${p}"
done
printf 'no-slash redirect: '; curl -s -o /dev/null -w '%{http_code} -> %{redirect_url}\n' "${SITE}/about"
Expect: / and /index.html → 200 text/html (the old page still
serves — the router did not break it); /about/ and /nope-not-a-page/ → 404
rather than the 403 they returned before, because Part 1 changed S3's answer and
Part 3 mapped it; /about → 301 → https://adr.smlcompany.ca/about/.
The body of those 404s is not the 404 page yet — 404.html is not in the bucket
until Part 8. The status is what Part 4 proves; the page is what Part 8.4
proves. Do not tick docs/06's "404 returns a 404 status" item here; tick it
after 8.4, when both halves are true at once.
Part 5 — Deploy the handler
5.1 Package
rm -f /tmp/intake.zip
(cd backend/intake && zip -q -X /tmp/intake.zip handler.mjs fields.mjs)
unzip -l /tmp/intake.zip
Expect: exactly two entries, handler.mjs and fields.mjs, ≈ 25.7 KB
uncompressed and ≈ 10.8 KB zipped [measured 2026-09-01]. Both at the zip root —
handler.mjs imports ./fields.mjs, so a nested directory breaks the import at
cold start.
5.2 Configuration first, code second
The handler calls requireEnv() at module scope and throws at cold start on
any of six missing variables — deliberately, so a misconfigured function cannot
accept a submission it will not store. Set the configuration before the code, so
there is never a moment when the new code runs against no environment.
ENV_JSON=$(node scripts/intake-env.mjs --table "$TABLE" --notify "$NOTIFY_TO" --from "$MAIL_FROM")
echo "$ENV_JSON" | python3 -m json.tool
Expect six variables. ⚠️ RESPONSE_TIME and NO_RETAINER_NOTICE are read
out of src/data/site.ts and must never be typed here. Both are published
commitments — the two-business-day response (§4, Q27) and the no-retainer notice
docs/01 §/contact/ requires, including its fourth clause about not itself
creating a conflict check, which a hand-typed copy inside the handler had
already dropped once. The notice also contains an en dash in "mediator–party",
which is exactly the character a retype loses. scripts/intake-env.mjs asserts
both and exits non-zero rather than emitting a softened commitment.
aws lambda update-function-configuration --function-name "$FN" \
--handler handler.handler --timeout 15 --memory-size 512 --environment "$ENV_JSON"
aws lambda wait function-updated --function-name "$FN"
aws lambda get-function-configuration --function-name "$FN" \
--query '{Handler:Handler,Timeout:Timeout,Memory:MemorySize,Vars:sort(keys(Environment.Variables)),State:State,Update:LastUpdateStatus}'
Expect: Handler: handler.handler — the file is handler.mjs and the export
is handler, where the hand-built function was index.handler; Timeout: 15;
Memory: 512; the six names sorted; State: Active; Update: Successful.
Timeout 10 → 15 s and memory 128 → 512 MB are changes, not restatements. The handler does a DynamoDB write and two SES sends per invocation, and 128 MB is where an AWS SDK cold start is slowest — more memory shortens the billed duration enough that it is usually the cheaper setting, not the dearer one.
5.3 Code
aws lambda update-function-code --function-name "$FN" --zip-file fileb:///tmp/intake.zip
aws lambda wait function-updated --function-name "$FN"
aws lambda get-function-configuration --function-name "$FN" \
--query '{CodeSize:CodeSize,Runtime:Runtime,Update:LastUpdateStatus,Modified:LastModified}'
Expect: CodeSize ≈ 10,800 (up from 1,527), Update: Successful.
⚠️ CodeSize is the ZIP, not the source. This line said "around 23,000",
which was 5.1's uncompressed figure applied to a different quantity — an
operator seeing 10819 against an expectation of 23,000 would reasonably
conclude the wrong artefact went up.
5.4 Prove it loads, without writing anything
cat > /tmp/probe-no-origin.json <<'JSON'
{"version":"2.0","routeKey":"POST /api/intake","rawPath":"/api/intake",
"headers":{"content-type":"application/x-www-form-urlencoded"},
"requestContext":{"http":{"method":"POST","path":"/api/intake","sourceIp":"127.0.0.1"}},
"body":"probe=1","isBase64Encoded":false}
JSON
aws lambda invoke --function-name "$FN" --cli-binary-format raw-in-base64-out \
--payload file:///tmp/probe-no-origin.json /tmp/probe-out.json
cat /tmp/probe-out.json; echo
Expect: the invoke prints "StatusCode": 200 with no FunctionError, and
/tmp/probe-out.json is
{"statusCode":403,"headers":{"Cache-Control":"no-store"},"body":""}
That 403 is the handler's Origin check refusing a request with neither Origin
nor Referer. It happens before any DynamoDB write and before any email,
which is what makes this probe safe to run against production.
Two other outcomes, and each has one cause:
"FunctionError": "Unhandled"withRuntime.ImportModuleErrorin the body — thenodejs24.xruntime does not provide the AWS SDK v3 clients this handler imports. Run 5.5, then repeat 5.4.- A body naming
intake handler: <NAME> is not set— that variable did not reach the function. Re-run 5.2 and read theVarslist.
5.5 Bundled variant — only if 5.4 said Runtime.ImportModuleError
Versions are resolved from the registry at run time rather than pinned in this
file: CLAUDE.md's rule is that a version is checked against the registry and
never recalled, and a literal here would be stale the week after it was written.
rm -rf /tmp/intake-bundle && mkdir -p /tmp/intake-bundle
cp backend/intake/handler.mjs backend/intake/fields.mjs /tmp/intake-bundle/
( cd /tmp/intake-bundle \
&& npm init -y > /dev/null \
&& npm install --omit=dev --no-audit --no-fund \
"@aws-sdk/client-dynamodb@$(npm view @aws-sdk/client-dynamodb version)" \
"@aws-sdk/client-sesv2@$(npm view @aws-sdk/client-sesv2 version)" )
rm -f /tmp/intake.zip
( cd /tmp/intake-bundle && zip -qr -X /tmp/intake.zip handler.mjs fields.mjs node_modules package.json )
unzip -l /tmp/intake.zip | tail -1
aws lambda update-function-code --function-name "$FN" --zip-file fileb:///tmp/intake.zip
aws lambda wait function-updated --function-name "$FN"
Expect a zip in the low single-digit megabytes, well under the 50 MB direct-
upload limit. If this path is taken, say so in the AGENTS.md Change Log and
add both packages to §7 — they become pins this project maintains, and R11
covers them from that moment.
Part 6 — Route the API at /api/intake
6.1 The invoke permission, before the route
⚠️ THE EXISTING PERMISSION DOES NOT COVER THE NEW ROUTE, AND THIS IS THE STEP
WHOSE OMISSION IS HARDEST TO DIAGNOSE. Part 0.3 records the one statement on the
function: SourceArn …/4tl0m5igkj/*/*/submissions. Add a route at
/api/intake without adding a permission for it and API Gateway is refused
permission to invoke — the caller gets a 500, the Lambda logs nothing at all
because it was never entered, and the only trace is an API Gateway metric.
aws lambda add-permission --function-name "$FN" \
--statement-id apigw-post-api-intake \
--action lambda:InvokeFunction \
--principal apigateway.amazonaws.com \
--source-arn "arn:aws:execute-api:${AWS_REGION}:${ACCT}:${API_ID}/*/POST/api/intake"
aws lambda get-policy --function-name "$FN" --query Policy --output text | python3 -m json.tool
Expect: two statements — the old one scoped to …/submissions and the new one
to …/POST/api/intake.
6.2 The route
export INTEG_ID=$(aws apigatewayv2 get-integrations --api-id "$API_ID" \
--query "Items[?IntegrationUri=='arn:aws:lambda:${AWS_REGION}:${ACCT}:function:${FN}'].IntegrationId | [0]" \
--output text)
echo "INTEG_ID=$INTEG_ID"
[ -n "$INTEG_ID" ] && [ "$INTEG_ID" != "None" ] || echo "STOP — no integration points at $FN"
aws apigatewayv2 create-route --api-id "$API_ID" \
--route-key 'POST /api/intake' --target "integrations/${INTEG_ID}" \
--query '{RouteId:RouteId,RouteKey:RouteKey,Target:Target}'
aws apigatewayv2 get-routes --api-id "$API_ID" --query 'Items[].{Key:RouteKey,Target:Target}'
Expect: INTEG_ID non-empty and not None; then two routes,
POST /submissions and POST /api/intake, both on the same integration. The
stage is $default with auto-deploy on, so the route is live immediately — there
is no deployment to create.
6.3 Throttling
⚠️ THIS IS NOT THE PER-IP LIMIT docs/05 ASKS FOR, AND IT CANNOT BE. That
spec says "Rate limit by source IP at API Gateway: 5 requests / 5 minutes".
API Gateway throttling is aggregate — per route and per stage — not per source
IP. Per-IP rate limiting needs AWS WAF with a rate-based rule on the
distribution, which is a paid service and therefore Pouya's decision, recorded on
docs/06's checklist as explicitly not a launch blocker. What is set here is
a real control over total volume; describing it as per-IP would be the Q22 shape
again — a documented control that does not exist.
aws apigatewayv2 update-stage --api-id "$API_ID" --stage-name '$default' \
--route-settings '{"POST /api/intake":{"ThrottlingRateLimit":1,"ThrottlingBurstLimit":5,"DetailedMetricsEnabled":true}}'
aws apigatewayv2 get-stage --api-id "$API_ID" --stage-name '$default' \
--query 'RouteSettings'
Expect: the POST /api/intake entry with rate 1/s, burst 5, detailed metrics
on. One per second sustained with a burst of five is far above any human filling
this form and far below anything worth paying for.
6.4 Retire the old route
Nothing posts to /submissions any more — the form posts to /api/intake and
the handler that answered the old contract is gone. An unused public route on the
same function is surface with no purpose.
export OLD_ROUTE=$(aws apigatewayv2 get-routes --api-id "$API_ID" \
--query "Items[?RouteKey=='POST /submissions'].RouteId | [0]" --output text)
echo "OLD_ROUTE=$OLD_ROUTE"
aws apigatewayv2 delete-route --api-id "$API_ID" --route-id "$OLD_ROUTE"
aws lambda remove-permission --function-name "$FN" \
--statement-id 70aeb597-e4d8-5141-8fb1-6b6190f7b2ab
aws apigatewayv2 get-routes --api-id "$API_ID" --query 'Items[].RouteKey'
aws lambda get-policy --function-name "$FN" --query Policy --output text | python3 -m json.tool
Expect: one route, POST /api/intake; one permission statement, scoped to it.
Confirm the statement id in the remove-permission call against the policy you
printed in 6.1 before running it — it is an id, not a name, and it will differ if
the permission was ever rebuilt.
The
execute-apihostname stays reachable, so/api/intakecan be reached without passing through CloudFront. That is not a new weakness and not a loose end:DisableExecuteApiEndpointcannot be turned on, because the distribution's own origin is that hostname. TheOrigincheck is a CSRF control rather than authentication either way, and 6.3's throttle applies at the stage, so it covers both paths.
Part 7 — End to end, with a real submission
7.1 The route answers
curl -si -X POST "${SITE}/api/intake" \
-H 'Origin: https://adr.smlcompany.ca' \
-H 'Content-Type: application/x-www-form-urlencoded' \
--data 'deploy-route-probe=1' | head -12
Expect: HTTP/2 303 and location: https://adr.smlcompany.ca/contact/could-not-send/.
The handler accepted the request, validated it, found an empty submission and
redirected to the failure page — before any write and any email. This is the
same probe scripts/deploy-local.sh runs at the end of every deploy.
⚠️ 404 IS AMBIGUOUS AND THE FIRST DRAFT OF THIS LINE WAS NOT. It means
either the /api/* behaviour is missing (Part 3) or the POST /api/intake route is missing or misspelled (Part 6.2) — two different repairs.
And because Part 3's custom error response is distribution-wide, a 404 from
API Gateway is served as the styled /404.html body, so the one string that would
have named the cause — API Gateway's {"message":"Not Found"} — is replaced before
you see it. Check the route first; it is one command:
aws apigatewayv2 get-routes --api-id "$API_ID" --query 'Items[].RouteKey'.
403 means the Origin header did not arrive — check that the behaviour uses
Managed-AllViewerExceptHostHeader, because a policy that drops Origin turns
every real submission into a 403. 500 means Part 6.1 was skipped.
7.2 A real submission, from the real form
Do this in a browser at ${SITE}/contact/ after Part 8, because the form
does not exist until the site is deployed. Fill it as an inquirer would, with
Your name: Cutover test <date> so the record is identifiable, and a real
address you can read.
Expect: the browser lands on /contact/received/, and two emails arrive — the
notification at $NOTIFY_TO, replyable to the address you entered, and the
confirmation at that address. Read the confirmation and check three things: the
response-time sentence, the four-clause no-retainer notice, and that the
field summary uses readable labels ("Your name", "Subject matter") rather than
field names.
Then find the record. The notification email prints submissionId <uuid> —
that value is the partition key, so it can be used directly:
export SUB_ID='<the uuid from the notification email>'
aws dynamodb get-item --table-name "$TABLE" \
--key "{\"submissionId\":{\"S\":\"${SUB_ID}\"}}" \
--query 'Item.{id:submissionId.S,at:submittedAt.S,ttl:ttl.N,name:name.S,area:practiceArea.S,sourceIp:sourceIp.S}'
curl -s https://checkip.amazonaws.com
⚠️ READ sourceIp AGAINST YOUR OWN ADDRESS — THIS IS THE ONLY PLACE THE PROXY
CHAIN GETS MEASURED, AND IT HAS ALREADY BEEN REASONED ABOUT WRONGLY TWICE. The
handler stores requestContext.http.sourceIp, which is the TCP peer. Behind the
/api/* behaviour that peer is a CloudFront edge, so the expectation is an
AWS address, not yours. Three outcomes and each has a different consequence:
what sourceIp holds |
what it means |
|---|---|
an AWS address (not the checkip value) |
As designed. The field records the CDN, so it cannot serve abuse investigation, and /legal/privacy/'s "your IP address" is inaccurate — fold it into the Q62 edit on the same page rather than leaving two wrong sentences there |
your address, matching checkip |
Better than expected, and worth knowing before anyone relies on it. Do not conclude it is trustworthy: verify it is not simply echoing a header by resubmitting with -H 'X-Forwarded-For: 8.8.8.8' and confirming 8.8.8.8 is not what lands. ⚠️ AND CORRECT /legal/privacy/ §What is collected, whose network-address paragraph says the address is "normally the network's own rather than your connection's" — wrong in this branch, and it understates what is held about the reader. This row carried no instruction at all until 2026-09-02, so two of these three outcomes had nothing reconciling the page with the measurement (adversarial-reviewer, round 2) |
8.8.8.8 after that resubmission |
Stop. The field is client-controlled and a record can be made to name an uninvolved third party. Revert to storing nothing rather than storing that |
⚠️ AN EARLIER REVISION OF THE HANDLER READ x-forwarded-for HERE, AND THAT WAS
THE THIRD OUTCOME. CloudFront appends the viewer address to a
client-supplied XFF rather than replacing it, so the leftmost entry is whatever
the client sent. The fix, if a usable value is wanted, is a custom origin
request policy on /api/* forwarding CloudFront-Viewer-Address, which
CloudFront generates and overwrites — not the managed
AllViewerAndCloudFrontHeaders, which forwards Host and would 403 every request
at API Gateway. That is an infrastructure change and it is deliberately not in
this runbook: measure first.
Expect the item, with ttl a 10-digit epoch-seconds value. Check it is 24
months out — read it, do not assume it:
python3 - <<PY
import datetime, os
ttl = int("${SUB_ID}" and os.popen(
'aws dynamodb get-item --table-name ${TABLE} '
'--key \'{"submissionId":{"S":"${SUB_ID}"}}\' --query Item.ttl.N --output text'
).read().strip())
d = datetime.datetime.fromtimestamp(ttl, datetime.timezone.utc)
now = datetime.datetime.now(datetime.timezone.utc)
months = (d.year - now.year) * 12 + (d.month - now.month)
print(f'ttl {ttl} -> {d.isoformat()} ({months} months from now)')
print('PASS' if months == 24 else 'FAIL — /legal/privacy/ promises 24 months')
PY
Expect: 24 months from now and PASS. /legal/privacy/ states the period
publicly, so a number that is not 24 is a false disclosure rather than a config
error.
⚠️ DELETE THE TEST RECORD when you are done with it — it is a real row in a table whose contents are governed by a published privacy policy.
aws dynamodb delete-item --table-name "$TABLE" --key "{\"submissionId\":{\"S\":\"${SUB_ID}\"}}"
7.3 The four pre-existing items
The table held 4 items before any of this [verified 2026-09-01 — describe-table ItemCount]. They were written by the hand-built handler, whose
code is in docs/reference/AWS-Hosting-Guide.md Part 8.3 and writes no ttl
attribute — so DynamoDB will never expire them and they are retained
indefinitely, against a policy that says otherwise.
This is Pouya's call and not a command in this runbook, because it depends on
something no agent can determine: whether those four are the guide's own smoke
test (Test User) or real inquiries that arrived through the old site. Read them
in the console, then either delete the test rows or write a ttl onto the real
ones. Tracked on docs/06's checklist.
Part 8 — The first production deploy
⚠️ DO NOT RUN 8.3 UNTIL POUYA HAS READ EVERY PAGE. docs/06's Content and
compliance group — the claims-auditor pass and his own page-by-page read — is
the gate on this part, and it is not a technical check.
8.1 Archive what is there
S3 versioning is Enabled, so pass 3's --delete is recoverable — but
docs/06's post-cutover list says to archive the old build rather than rely on
that, and a local copy costs one command.
mkdir -p _archive/pre-cutover-$(date +%F)
aws s3 sync "s3://${BUCKET}" "_archive/pre-cutover-$(date +%F)/" --no-progress
ls -la "_archive/pre-cutover-$(date +%F)/"
find "_archive/pre-cutover-$(date +%F)" -type f | wc -l
Expect: 3 files — index.html at 2,206,032 bytes and two PNGs under
assets/. Do not commit them; docs/06 wants them kept, not versioned.
8.2 Dry run — the three passes, in order, writing nothing
npm run check && npm run build && npm run check:claims && npm run og:proof && npm run check:intake
echo "gates exit=$?"
npm run lighthouse; echo "lighthouse exit=$?"
Expect: every one exit 0. lighthouse is local-only and cannot run in CI; it
is a keyboard gate and this is the keyboard.
aws s3 sync ./dist "s3://${BUCKET}" --dryrun \
--exclude "*" --include "_astro/*" --include "fonts/*" \
--cache-control "public, max-age=31536000, immutable" --no-progress | tail -5
aws s3 sync ./dist "s3://${BUCKET}" --dryrun \
--exclude "*" --include "*.avif" --include "*.webp" --include "*.jpg" \
--include "*.png" --include "*.svg" \
--cache-control "public, max-age=604800" --no-progress | tail -5
aws s3 sync ./dist "s3://${BUCKET}" --dryrun \
--exclude "_astro/*" --exclude "fonts/*" \
--cache-control "public, max-age=0, must-revalidate" --delete --no-progress | tail -8
Expect, on the bucket state Part 0.3 records — measured 2026-09-01, so these
are counts rather than shapes: 51 (dryrun) upload: lines in pass 1, 56
in pass 2, and 52 lines in pass 3 of which two are
(dryrun) delete: — assets/sml-logo-full.png and assets/sml-logo-mark.png,
the old site's two logos. 404.html appears in pass 3. index.html is not
deleted; it is overwritten.
⚠️ PASS 2's DRY RUN IS A SUPERSET OF WHAT PASS 2 ACTUALLY UPLOADS, AND THE REAL
RUN WILL PRINT FEWER LINES. 34 of those 56 are files under _astro/, which
pass 1 uploads first. Each --dryrun above is compared against the bucket as it
is now, so it cannot see the uploads the preceding pass would have made; in the
real sequential run aws s3 sync skips them as already in sync. That skipping is
exactly what preserves pass 1's immutable header on those files rather than
overwriting it with pass 2's week-long one — docs/06 §Cache policy says the same
thing about pass 3, and this is why the pass order is load-bearing. Do not
reorder the passes to make this output tidier, and do not read the smaller real
counts as a failed upload.
⚠️ Check the largest file in ./dist while you are here. The deploy user has
no s3:AbortMultipartUpload, which is safe only while nothing approaches
aws s3 sync's 8 MB multipart threshold (docs/06).
find ./dist -type f -exec stat -f '%z %N' {} + | sort -rn | head -3
Expect the largest well under 8,388,608 bytes — it was 171,876 on
2026-09-01. GNU find -printf is not available here; this is the BSD/macOS form,
and writing it as find -printf … || find -exec … would not have fallen back,
because the exit status of that pipeline belongs to head.
8.3 The deploy — on Pouya's word only
⚠️ NOT AS user/pouya. The script refuses that identity and §10 is why: the
account is shared with unrelated production systems. Use the adr-sml-deploy
key, which §7 records as provisioned.
AWS_ACCESS_KEY_ID=… AWS_SECRET_ACCESS_KEY=… \
AWS_REGION="$AWS_REGION" S3_BUCKET="$BUCKET" CLOUDFRONT_DISTRIBUTION_ID="$DIST_ID" \
npm run deploy
The script runs npm run check, the build, check:claims, the three passes, the
/* invalidation, and then the intake route check. Read its last line. It
warns rather than fails on a bad intake route, because by then the site is
already published and failing the job would not un-publish it.
git tag "v$(date +%Y).1" && git push --tags
8.4 Verify the site, not the deploy
for p in / /about/ /mediation/ /arbitration/ /med-arb/ /practice/ /practice/construction/ \
/process/ /for-parties/ /fees/ /insights/ /contact/ /bio/ /legal/privacy/ \
/legal/terms/ /robots.txt /sitemap-index.xml /pouya-lajevardi-bio.pdf; do
printf '%-32s ' "$p"
curl -s -o /dev/null -w 'status=%{http_code} bytes=%{size_download}\n' "${SITE}${p}"
done
printf '%-32s ' "404 body"; curl -s "${SITE}/nope-not-a-page/" -o /tmp/404.html -w 'status=%{http_code}\n'
grep -c 'That page is not here' /tmp/404.html
printf '%-32s ' "cache header on HTML"; curl -sI "${SITE}/about/" | grep -i '^cache-control'
printf '%-32s ' "cache header on asset"
curl -s "${SITE}/about/" | grep -o '/_astro/[^"]*\.css' | head -1 \
| xargs -I{} curl -sI "${SITE}{}" | grep -i '^cache-control'
Expect: every page 200 with real bytes; the 404 path 404 and the phrase
That page is not here found once; max-age=0, must-revalidate on HTML;
max-age=31536000, immutable on the hashed CSS. Only now is docs/06's "404
returns a 404 status" item true on both halves.
Part 9 — Rollback, per part
Each of these is independent. None of them needs the others undone first.
9.1 Part 1 — aws s3api put-bucket-policy --bucket "$BUCKET" --policy "file:///tmp/bucket-policy.before.json".
Missing keys go back to 403 and the 404 mapping stops firing; nothing else changes.
9.2 Parts 2–3 — re-run configure.mjs is not a rollback; it is idempotent
forward-only. To undo, get-distribution-config, remove the
FunctionAssociations entry / the 404 custom error response / the /api/*
behaviour and the intake-api origin, and update-distribution --if-match. Then
aws cloudfront delete-function --name adr-sml-router --if-match <etag>, which
fails while the function is still associated — that ordering is a feature.
9.3 Part 5 — the previous code is not recoverable from Lambda: the hand-built
function was edited in the console and its 1,527 bytes exist only in
docs/reference/AWS-Hosting-Guide.md Part 8.3. That is the rollback source, and
it answers a different contract (JSON in, {"ok":true} out) that the current
/contact/ form does not speak. Forward is the only real fix here — which is
worth knowing before 5.3, not after.
9.4 Part 6 — delete-route the new route, remove-permission the new
statement, and re-add the /submissions route and its permission if anything
still needs it. Nothing does.
9.5 Part 8 — S3 versioning is Enabled. Restore the previous version of
index.html, aws s3 rm what the new build added, invalidate /*. Or
git revert and re-deploy, which is cleaner and is what the tag in 8.3 is for.
Part 10 — Q60: prove TTL deletes, not just that it is enabled
AGENTS.md §7 records TimeToLiveStatus: ENABLED on attribute ttl
[re-verified 2026-09-01], and that proves the setting and not the behaviour.
/legal/privacy/ tells the public that a record is "deleted automatically by the
database rather than by someone remembering to do it" — an assertion about the
mechanism. Q60 closes when a record has been watched to disappear, and not
before. §12 R19 keeps it surfacing until then.
10.1 Write one record with a near-future ttl
Not through the form: the handler writes a 24-month ttl and waiting two years
is not a test. Write it directly, with a shape that cannot be mistaken for an
inquiry.
export TTL_ID="q60-ttl-probe-$(date +%Y%m%dT%H%M%SZ)"
export TTL_AT=$(python3 -c 'import time; print(int(time.time()) + 300)')
python3 -c "import datetime,os; print('ttl', os.environ['TTL_AT'], '->', datetime.datetime.fromtimestamp(int(os.environ['TTL_AT']), datetime.timezone.utc).isoformat())"
aws dynamodb put-item --table-name "$TABLE" --item "$(cat <<JSON
{"submissionId":{"S":"${TTL_ID}"},
"submittedAt":{"S":"$(date -u +%Y-%m-%dT%H:%M:%SZ)"},
"ttl":{"N":"${TTL_AT}"},
"note":{"S":"AGENTS.md Q60 — TTL behaviour probe, no personal information, safe to delete"}}
JSON
)"
aws dynamodb get-item --table-name "$TABLE" --key "{\"submissionId\":{\"S\":\"${TTL_ID}\"}}" \
--query 'Item.{id:submissionId.S,ttl:ttl.N,note:note.S}'
echo "$TTL_ID" > ~/q60-ttl-probe-id.txt
echo "$TTL_AT" >> ~/q60-ttl-probe-id.txt
Expect: the item, with ttl five minutes in the future. The id is written to
a file because the check below happens on a different day and this is the only
thing that connects the two.
10.2 The check, and it is not five minutes later
⚠️ DYNAMODB'S TTL SWEEPER IS NOT PROMPT AND AWS DOES NOT PROMISE THAT IT IS.
Deletion typically happens within a couple of days of the timestamp passing, and
up to 48 hours or more is normal and not a fault. A check run at ttl + 5 minutes that finds the item present has measured nothing — it is the same
mistake as reading a value before the transition finishes. So:
- Earliest useful check: 48 hours after
TTL_AT. - Do not conclude a failure before 7 days.
export TTL_ID=$(sed -n 1p ~/q60-ttl-probe-id.txt)
export TTL_AT=$(sed -n 2p ~/q60-ttl-probe-id.txt)
python3 -c "import time,os; d=time.time()-int(os.environ['TTL_AT']); print(f'{d/3600:.1f} hours since the ttl passed')"
aws dynamodb get-item --table-name "$TABLE" \
--key "{\"submissionId\":{\"S\":\"${TTL_ID}\"}}" --output json
echo "get-item exit=$?"
Read both the output and the exit status. get-item on a missing key exits
0 with an empty response body — it is not an error, and a script that only
checked the exit status would report the record present and absent identically.
| result | meaning |
|---|---|
{"Item": {…}}, under 48 h since TTL_AT |
inconclusive. Wait. |
{} (empty), any time after TTL_AT |
Q60 closes. The mechanism works. |
{"Item": {…}}, more than 7 days after |
Q60 fails. TTL is enabled and not deleting. /legal/privacy/ is then making a claim the database does not honour, and the page has to change or the mechanism does. |
10.3 Record it
On the {} result, and only then:
- Stamp §7's
Intake tablerow: the behaviour is observed, with the date and the elapsed time, alongside theENABLEDreading it already carries. - Close Q60 in §9, quoting the two timestamps.
- Retire R19 from §12 — its whole purpose was to keep this surfacing.
- Tick the TTL item on
docs/06's checklist. It has two halves and this is the second:ENABLEDproved the setting, this proves the behaviour. - Delete the probe row if it somehow survives, and delete
~/q60-ttl-probe-id.txt.
If it fails, that is a published-disclosure defect, not a backlog item:
src/pages/legal/privacy.astro carries the matching TODO(pouya) and the page
asserts the mechanism in as many words.